A Schema Registry stores and versions schemas for Kafka message keys/values (commonly Avro, Protobuf, or JSON Schema).
Problem without it
Producers and consumers share bytes on a topic. If the producer adds a field and the consumer is not ready, or removes a field the consumer needs, you get poison messages and runtime breakage.
What Schema Registry does
- Stores schemas under a subject (often
<topic>-value/<topic>-key) - Assigns a numeric schema id
- Producer serializers register/lookup schema and typically embed the schema id in the message payload
- Consumer deserializers fetch schema by id and decode safely
- Enforces compatibility rules on schema evolution
Producer --schema--> Schema Registry (id=42) Producer --bytes(+id)--> Kafka topic Consumer --id 42--> Schema Registry --> decode with schema
Why data engineers care
Kafka moves data between many teams. Schema Registry is the contract layer: evolve carefully, fail registration when incompatible, avoid silent corruption.
Common products
Confluent Schema Registry is the most cited. Similar ideas exist in other stacks (AWS Glue Schema Registry, Apicurio, etc.).
Interview tip: "Schema Registry = versioned contracts for Kafka payloads + compatibility checks during evolution."