Skip to content
LakeBench
ProblemsCommunityPricing
Sign inStart practicing
Back
  1. Home
  2. Interview prep
  3. Hashing vs encryption vs tokenization

Data quality · Governance, Privacy & Lineage

Hashing vs encryption vs tokenization

Mediumdata-quality-30
cryptographyhashingencryptiontokenization

Question

What is the difference between hashing, encryption and tokenization for sensitive data?

Solution

Hashing is a one-way mathematical function that cannot be reversed, encryption is a two-way reversible algorithm requiring cryptographic keys, and tokenization replaces sensitive data with non-sensitive surrogate tokens stored in a secure lookup vault. Choosing the right technique depends on whether downstream analytics requires joining datasets, recovering original values, or minimizing breach exposure.

Comparing masking and security primitives

Each cryptographic primitive serves distinct operational requirements:

  • Hashing: Produces a deterministic, fixed-length digest from an input string using algorithms like SHA-256. Because identical inputs yield identical outputs, hashed fields support joins across datasets without exposing raw identifiers. However, low-entropy values like phone numbers or postal codes are vulnerable to rainbow table and brute-force dictionary attacks unless secured with cryptographic salts and secret peppers.
  • Encryption: A reversible mathematical operation using symmetric algorithms like AES-256 or asymmetric public-private keypairs. Anyone holding the authorized decryption key can recover original cleartext. Encryption suits operational systems that must retrieve the original value later, but running analytical joins or group-by aggregations on ciphertext requires specialized deterministic encryption.
  • Tokenization: Replaces the sensitive value with an arbitrary, randomly generated token that has no mathematical relationship to the original plaintext. The mapping between the token and raw data is stored in an isolated, access-controlled vault. If the data warehouse is breached, attackers cannot derive original values from the tokens.
Technique     Reversible?   Joinable?   Vault Needed?  Primary Use Case
Hashing       No            Yes         No             Anonymized cross-table joins
Encryption    Yes (Key)     Difficult   No (Keys only) Reversible storage for ops
Tokenization  Yes (Vault)   Yes         Yes            Payment cards, national IDs

Selecting an approach requires evaluating analytical and recovery needs:

  • Choose hashing when you need to join user behavior across tables without needing to recover the underlying personal identifier.
  • Choose encryption when authorized services must regularly decrypt and restore the original value.
  • Choose tokenization when handling high-risk fields like payment cards where storing ciphertext inside the analytics lakehouse poses unacceptable compliance risk.

Downstream analytical trade-offs

Applying these primitives directly impacts warehouse performance and analytical flexibility. Storing irreversible hashes preserves join capability while eliminating privacy risk, whereas managing encryption keys introduces decryption overhead on analytical queries.

PreviousNext