Python data engineering interview problem. Difficulty: intermediate. Pattern: Graphs. About 15 minutes. Part of the Pro drill bank.
Production ticket: A scheduler can't just run tasks in the order they were declared - it has to respect the dependency graph. Treat this as a production helper: match the contracted return shape, including empty and duplicate inputs.
A DAG config lists each task and the tasks it depends on: {"extract": [], "transform": ["extract"], "load": ["transform"]}. The scheduler must run every dependency before the task that needs it. Write topo_order(deps: dict[str, list[str]]) -> list[str] returning one valid execution order. When more than one task is ready at the same time, run them in the order they first appear as keys in deps. Keep the prints at the bottom. The grader reads stdout.
Input: topo_order({"extract": [], "transform": ["extract"], "load": ["transform"], "notify": ["load"]}) Output: ['extract', 'transform', 'load', 'notify'] This is the expected return value for the given call.
Topics: orchestration, dag, topological-sort.
More Python interview questions · All interview problems · Learn data engineering
Production ticket: A scheduler can't just run tasks in the order they were declared - it has to respect the dependency graph.
A DAG config lists each task and the tasks it depends on: `{"extract": [], "transform": ["extract"], "load": ["transform"]}`. The scheduler must run every dependency before the task that needs it. Write `topo_order(deps: dict[str, list[str]]) -> list[str]` returning one valid execution order. When more than one task is ready at the same time, run them in the order they first appear as keys in `deps`. Keep the prints at the bottom. The grader reads stdout.