In the Hadoop Distributed File System (HDFS), the NameNode manages the file system namespace and metadata in memory, while DataNodes store the physical data blocks on local server disks with three-way replication. Early Hadoop architectures suffered from the NameNode being a single point of failure, which was later resolved using active-standby High Availability configurations coordinated by ZooKeeper. Technical interviewers continue asking about Hadoop because major enterprise sectors still maintain on-premise clusters, and core distributed storage concepts like data locality and block sizing directly inform modern engines like Apache Spark.
Responsibilities of NameNode and DataNodes
HDFS separates metadata orchestration from physical block storage:
- The NameNode: Holds the entire file system directory tree, permissions, and file-to-block mappings in JVM heap memory. Because every file, directory, and block requires approximately 150 bytes of memory, storing millions of small files exhausts NameNode memory, which is the classic Hadoop small-file problem.
- The DataNodes: Worker servers that store data blocks (typically 128 MB or 256 MB in size) on local hard drives. DataNodes read and write blocks on request and send periodic block-report heartbeats back to the NameNode to confirm disk health. By default, each block is replicated across three distinct nodes and racks to tolerate hardware failure.
Solving the single point of failure
In Hadoop 1.x, a cluster had only one NameNode. If that server experienced hardware failure, the entire file system became unavailable until manual recovery was performed.
Hadoop 2.x and 3.x introduced High Availability (HA) architectures:
Client Requests ---> Active NameNode <--- Quorum Journal Nodes ---> Standby NameNode
| |
+---------------- ZooKeeper Failover ------------+An Active NameNode and Standby NameNode synchronize edits through Quorum Journal Nodes, while ZooKeeper monitors heartbeats to execute automatic failovers if the active node crashes.
Why interviewers still evaluate Hadoop knowledge
Interviewers ask about Hadoop for two pragmatic reasons:
First, massive legacy enterprises in banking, telecommunications, and healthcare continue operating petabyte-scale on-premise Hadoop clusters due to data residency regulations and infrastructure longevity.
Second, the foundational concepts of HDFS directly underpin cloud data engineering. Understanding why HDFS chose 128 MB blocks explains why Apache Spark partitions data at 128 MB boundaries, and understanding data locality clarifies why distributed query planners minimize network shuffle operations when processing petabytes of data.