IAM (Identity and Access Management) controls *who* can do *what* on which resources. Users, roles/service accounts, groups, and policies grant or deny actions like s3:GetObject or bigquery.jobs.create.
Human user / CI / Spark job
|
assumes ROLE / service account
|
POLICY allows: read lake bucket, write curated bucketCore ideas
- Identity: user, role, service account
- Policy: allow/deny statements on actions + resources
- Least privilege: grant only what the job needs
- Prefer roles for apps over long-lived user keys
DE failure mode
An ETL role with * on all buckets leaks PII if the job is compromised. Separate raw vs curated permissions.
Interview tip: "IAM = who can do what." Say least privilege + roles for pipelines, and never embed permanent admin keys in code.