A migration like this is mostly an inventory and people problem, and less about copying bytes. Plan it in phases, keep old and new running side by side for a while, and move consumers in waves.
Phase 1: inventory
List all tables (size, last read, owner), all jobs (Hive, Spark, Oozie or other schedulers), all consumers (dashboards, extracts, ML jobs), and the lineage between them. Often 30 to 40 percent of tables are unused, so use the access logs to find out, and do not migrate them. That cuts scope and cost.
Phase 2: choose the target and prioritise
Pick the target platform on workload and skills (BigQuery, Snowflake or Databricks), and decide the landing design (formats, naming, security model). Rank the work by business value and difficulty, and start with a vertical slice: one domain, end to end, from ingestion to dashboard. It proves the approach and teaches the team.
Phase 3: move data
Bulk copy the history (transfer appliances or network copies from HDFS to object storage), then keep it in sync with incremental loads or CDC until cutover. Convert to open columnar formats if not already (Parquet or ORC).
Phase 4: translate the code
Hive SQL and UDFs need rewriting to the new dialect. Use automated translation tools where they exist, and expect to fix a share of queries by hand: date functions, window details, null and cast behaviour. Orchestration moves too (Oozie to Airflow or Workflows). Treat each job as a small project with tests.
Phase 5: run in parallel and reconcile
Run old and new pipelines on the same inputs for a period, and compare outputs: row counts, aggregates by partition, and a sample of exact rows. Fix differences before consumers switch. This is the step that builds trust.
Phase 6: cut over in waves
Move consumers group by group, with a rollback plan. Freeze changes to the old system during each wave. After a stable period, shut down the old jobs and then the cluster.
Risks and extras
- Cost model change: from fixed cluster cost to usage-based. Set budgets, tagging and guardrails before users arrive.
- Governance: map Ranger or Sentry permissions to the new access model.
- People: training, and a support channel for the migration period.
- Do not migrate everything "as is" forever. A rewrite of the worst jobs is a good use of the move.