Securing a pipeline means protecting data in transit, at rest, and in use, plus locking down who can run jobs and read outputs.
Sources --TLS--> Ingestion --IAM--> Lake/WH --RBAC--> Consumers
^
secrets from vault (not hardcoded)Checklist
1. Encryption: TLS in transit; KMS encryption at rest on buckets/tables. 2. Secrets: API keys/DB passwords in Secret Manager / Vault, rotated. 3. Identity: workload IAM roles, not shared long-lived keys in code. 4. Access control: least privilege on raw vs gold; column/row security for PII. 5. Network: private connectivity where required (VPC, PrivateLink). 6. Audit: who read PII, who changed DAG code, who ran backfills. 7. Data handling: mask/tokenizeize in lower environments; careful logs (no dumping card numbers).
Common fresher mistake
Embedding cloud keys in GitHub Actions plaintext or printing PII in Spark logs.
Interview tip: Walk encryption → secrets → IAM/RBAC → audit. Mention PII separation between raw and serving layers.