Databricks medallion architecture is a recommended, not required, way to improve data progressively: bronze preserves incoming source data, silver validates and refines it, and gold prepares useful data products for consumers. The best implementations make each layer’s purpose clear, preserve enough detail to rebuild downstream data, and align pipelines and governance with actual workload needs.
What the three layers are for
Bronze, silver, and gold are stages of increasing data quality and consumer readiness, not simply three storage buckets. Databricks says following the pattern is a recommended best practice, not a requirement. Use layer boundaries where they make traceability, quality control, reuse, or publishing easier to manage.
As an Amazon Associate I earn from qualifying purchases.
| Layer | Purpose | Typical contents |
|---|---|---|
| Bronze | Preserve source data for traceability and downstream rebuilding. | Incrementally appended raw records with source and provenance metadata where useful. |
| Silver | Validate, refine, and integrate data into reusable representations. | Cleaned, typed, deduplicated, and often joined record-level data. |
| Gold | Serve defined business, analytical, or operational needs. | Dimensional models, metrics, aggregates, summaries, and consumer-ready products. |
Databricks’ medallion architecture documentation describes the progression from raw to validated and enriched data.
Free tools Windows power users keep installed
One-click scans. No signup required.
How to design bronze for replay and traceability
Keep bronze faithful to what arrived and apply minimal cleanup or validation. Its value is that downstream layers can be rebuilt when transformation logic changes, records need reprocessing, or a source issue must be investigated. Preserve useful source and provenance metadata rather than discarding context at ingestion.
#1 Best Overall
- Ingest incrementally where the source and workload support it.
- Retain most fields in flexible types such as strings, VARIANT, or binary when unexpected schema changes could otherwise break ingestion.
- Keep raw data access controlled and define its retention and lifecycle rather than treating raw as unrestricted or permanent.
- Use Unity Catalog managed tables as the standard lakehouse storage choice; Databricks also recommends Unity Catalog volumes for landing zones and raw unstructured data. External tables can be appropriate when data must remain at specific storage paths.
These storage recommendations are in Databricks’ medallion guidance and its current lakehouse design guidance.
How to make silver the trusted reusable layer
Build silver from bronze or existing silver tables. It is the right place to handle schema enforcement or evolution, nulls, duplicate records, late or out-of-order events, type casting, joins, and explicit data-quality rules. Maintain at least one validated, non-aggregated representation of each record so downstream users are not limited to precomputed summaries. Aggregations can live in silver when a real downstream need justifies them, but they typically belong in gold.
- For most append-only sources, read from bronze rather than writing ingestion directly into silver; schema changes or corrupt records can otherwise make ingestion fragile.
- Prefer streaming reads for most append-only inputs. Batch reads can suit small datasets, such as small dimensions.
- Keep the detail needed for analysis and machine learning, and document transformation logic, quality rules, and freshness expectations.
- Make quality rules stricter as data moves toward consumers, while checking ingestion too so source problems are visible.
Databricks’ medallion guidance recommends this validated, reusable role for silver.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #2
How to shape gold around consumers
Gold should answer a defined consumer need rather than become another raw-data store. Model data products for business intelligence, reporting, machine learning, or operational use; these may include dimensional models, shared metrics, aggregates, and summaries. Apply protections such as anonymization, row-level access, or column masking where the data and audience require them.
Decide who owns publication. A centralized, domain-based, or hybrid approach can work, provided teams know which products are shared and under what policies. In a hub-and-spoke model, Databricks recommends a shared hub for organization-wide data, domain-specific ingestion and curation, explicit publishing policies, and Unity Catalog catalogs that distinguish hub and domain assets. See Databricks’ architecture guidance.
Choose pipeline primitives by transformation shape
Databricks’ Lakeflow guidance distinguishes incremental row-level work from transformations that benefit from incremental refresh of joins or aggregations. Separate ingestion and transformation pipelines when practical so that a downstream failure does not unnecessarily stop data from landing in bronze, and so each can be scheduled, monitored, and troubleshot independently.
| Workload | Useful starting point | Why |
|---|---|---|
| Raw ingestion and incremental row-level filtering, cleaning, or parsing | Streaming tables | Fits incremental ingestion and row-oriented transformations. |
| Enrichment joins or complex aggregations that benefit from incremental refresh | Materialized views | Can refresh derived results incrementally and precompute gold summaries. |
The choice depends on workload semantics and current feature support. Verify the current Lakeflow documentation before specifying syntax or relying on release-specific capabilities.
Build quality and governance into every layer
Use checks throughout the flow: surface ingestion problems in bronze, enforce stronger validation in silver, and publish appropriately governed products in gold. Databricks identifies constraints, expectations, primary- and foreign-key metadata, and Lakehouse Monitoring as quality-related capabilities. Informational primary- and foreign-key metadata should not be mistaken for enforced constraints.
Use Unity Catalog for discovery and lineage, and organize catalogs and schemas according to the organization’s governance and ownership model. Databricks recommends managed tables across bronze, silver, and gold; plan governance rather than allowing unmanaged table sprawl or skipping checks to meet a delivery deadline. See Databricks’ lakehouse design guidance and Unity Catalog best practices.
Rank #4
Choose boundaries for the workload, not by rule
Not every team needs the same pipeline or governance arrangement. Decide based on the actual constraints:
- Latency and ingestion: choose batch, streaming, or change data capture according to source behavior and freshness needs.
- Governance ownership: determine whether data stewardship is centralized, domain-based, or hybrid.
- Consumer needs: preserve detailed, reusable silver data where consumers need it; publish gold marts or aggregates for well-defined use cases.
- Storage control: use managed tables as the recommended default, but consider external tables where fixed storage paths are required.
- Operational boundaries: keep ingestion and transformation together only when that is operationally simpler; separate them when independent scheduling and failure isolation matter.
The medallion pattern is useful when progressive quality, traceability, and clear publishing boundaries solve real problems. It does not require every dataset to have three physically distinct systems or every team to adopt identical layer boundaries.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




