Blog

What is medallion architecture?

October 5, 2026 16 min. read
What is medallion architecture?

Medallion architecture is a data design pattern that organizes data in a lakehouse into three progressive layers, Bronze, Silver, and Gold, with each layer representing a step up in quality and structure from the one before it. Data lands in its rawest form, moves through validation and cleansing, and finally reaches business consumers as trusted, analytics-ready, and AI-ready information. The pattern was popularized by Databricks and is closely associated with the Delta Lake ecosystem. It is sometimes called multi-hop architecture because data effectively hops from one layer to the next as it gets refined.

The pattern also gives engineers, analysts, and increasingly AI systems a shared, predictable place to work at the right level of trust. A data engineer troubleshooting an ingestion failure has no reason to touch the same table an executive dashboard is querying, and separating those concerns into distinct layers is most of what medallion architecture actually does.

The three layers of medallion architecture

Each layer in a medallion architecture takes on a specific job, moving data from raw and unexamined to structured and business ready as it progresses from Bronze to Silver to Gold.

LayerInput data typeTransformation appliedData consumersQuality levelExample use case
BronzeRaw, unprocessed data from source systems: files, streams, APIs, database extractsNone. Ingested as is, typically append-onlyData engineers; audit and replay processesUnvalidated; preserved exactly as receivedStoring raw transaction logs or clickstream events for replay and audit
SilverBronze data, plus reference and lookup dataValidation, deduplication, standardization, schema enforcementData analysts, downstream pipelines, data scientistsValidated and conformed; passed defined data quality rulesCleansed customer records with standardized addresses and deduplicated identities
GoldSilver data, aggregated and modeledAggregation, dimensional modeling, business logic, certification against a trust thresholdBusiness users, BI dashboards, ML models, AI agentsCertified as production readyDaily revenue aggregates powering an executive dashboard

Bronze layer: Raw data ingestion

The Bronze layer is the raw data landing zone, where data arrives from source systems without transformation. Nothing about the incoming records changes at this stage. Files, event streams, and database extracts are appended largely as they arrive, which preserves the original data for replay and audit long after the source system has moved on. That immutability matters more than it sounds like it should, since a pipeline bug three steps downstream is far easier to diagnose when the exact record that arrived on day one is still sitting there to compare against.

Bronze is also, somewhat counterintuitively, where profiling tends to start. Automated profiling examines the structure and content of incoming data, its distributions, its outliers, its missing values, without altering a single record, and what it turns up becomes the basis for the validation rules applied one layer later.

Silver layer: Cleansed and conformed data

The Silver layer is where validation, deduplication, schema enforcement, and filtering happen, and it is fair to call it the data quality layer of a medallion architecture. Records that satisfy the business’s defined rules move forward, and records that do not can be gated, quarantined, or blocked outright, which is a meaningfully different posture than checking quality after the fact and cleaning up whatever slipped through. The section on data quality further down goes into how that gating actually works in practice.

Gold layer: Business-ready data

The Gold layer holds aggregated, modeled data built for direct consumption: dimensional models in star or snowflake schema, metrics rolled up for a BI tool, features ready for a model. This is the layer where data stops being an engineering concern and starts being a business asset, and increasingly it is also where a dataset gets certified as production ready before anything, human or automated, is allowed to act on it.

Benefits of medallion architecture

Organizing data into Bronze, Silver, and Gold layers pays off in several concrete ways:

  • Data quality improves progressively, since each layer enforces a higher standard than the last, so problems get caught closer to their source, so problems get caught in the pipeline rather than in a business report (where the person who spots them is multiple handoffs away from the team where they originated)
  • Auditability improves, because Bronze preserves the raw data and each subsequent layer is a deliberate transformation of what came before. That structure makes data lineage far easier to trace: following a number back to its source means following the layers, rather than piecing it together from code, job logs, and whoever still remembers how the pipeline was built. 
  • The pattern scales cleanly, since storage and compute can grow independently for each layer, so a spike in raw ingestion does not force a rebuild of the Gold layer reporting tables that business users depend on.
  • Concerns stay separated: engineers own Bronze and the pipelines that move data forward, analysts and stewards own Silver’s rules, and business teams own what Gold surfaces, which means fewer people need write access to more of the pipeline.
  • Streaming and batch coexist, since the same three-layer structure accommodates a nightly batch load and a continuous event stream without forcing one approach to mimic the other.
  • Data gets reused rather than recomputed. Once a dataset is cleaned in Silver, every downstream Gold table and every analyst querying that Silver table benefits from that single pass of work, rather than each consumer reimplementing the same cleansing logic.

Medallion architecture best practices

Getting a medallion architecture right has less to do with the layer names and more to do with the discipline applied at each transition:

  • Enforce schema from the start, rather than treating structure as something to clean up after ingestion. A schema defined at Bronze, even loosely, gives every later transformation something stable to check against.
  • Keep Bronze immutable. Resist the temptation to patch or delete raw records once they have landed. If a source system sends bad data, that is useful information, and overwriting it destroys the audit trail the whole pattern is built to preserve.
  • Gate promotions between layers rather than only monitoring what has already landed. The Bronze-to-Silver boundary is the natural point to validate records with Data Quality Gates before they are allowed into Silver, so Silver holds validated records by construction instead of by hope.
  • Define ownership per layer, and treat it as a data governance question tied to specific critical data elements rather than assigned uniformly per table. Not all data carries equal weight. A fraction of it drives regulatory submissions, risk calculations, and customer-facing decisions, and that fraction is where governance effort belongs first.
  • Avoid promoting data to Gold prematurely. A dataset that has not cleared validation has no business feeding a dashboard, no matter how urgently someone wants the number.

Together, these practices turn medallion architecture from a storage convention into an operational discipline, one that builds trust into the pipeline itself rather than checking for it after the data has already reached its destination.

A concrete example: How medallion architecture works in practice

Picture a retail bank processing card and transfer transactions across its business. In the Bronze layer, raw transaction events land continuously from core banking systems and card processors, stored exactly as received: the same fields, the same timestamps, the same currency codes, whether or not those codes turn out to be valid.

In the Silver layer, that raw feed goes through the checks that turn it into something usable. Transactions get validated against reference data, so a currency code or merchant category code that does not match the official list gets flagged rather than silently accepted, and records missing a mandatory field, an account ID or a transaction currency, are quarantined instead of passed forward. Duplicate submissions, the kind that happen when a request times out and the client retries, get detected and collapsed into a single record, and customer records tied to those transactions are matched and merged wherever the same underlying customer shows up across more than one source system, so the bank is not left working from several partial views of the same person.

By the time that data reaches Gold, it has been aggregated into daily summaries: spend by category, balances, exposure by risk segment, all certified as production ready before they feed a regulatory report or a credit risk model. That last step, certification, is what separates a Gold table someone happened to build carefully from a Gold table the organization can actually stand behind when a regulator or a model asks where a number came from.

Data quality and trust in the medallion architecture: What happens at each layer

Most explanations of medallion architecture mention data quality as a benefit without saying what enforces it, and that gap is worth closing, because the mechanics differ meaningfully from layer to layer.

Data quality work happens on two planes that both need to be present: checks on data already at rest in a layer, and checks on data in motion as it moves between layers. A dataset can pass every at-rest check on a schedule and still let a bad batch through in between checks, and a pipeline can validate every record in flight and still accumulate drift that only becomes visible once enough of it has landed, so relying on only one of the two leaves a real gap.

At Bronze, the work is almost entirely at rest and entirely non-destructive. Profiling analyzes structure and content the moment data arrives, surfacing missing values, duplicates, outliers, and the statistical patterns that show where defects concentrate, all without touching the raw records themselves. This is also where terms get assigned: a column gets recognized and labeled as an email address, a national ID, or a payment reference, and the subset of data that actually matters to the business, the critical data elements that drive regulatory submissions or risk calculations, gets flagged for closer attention. Bronze stays untouched through all of this. The goal here is discovery, not enforcement.

Silver is where enforcement happens. Validation, deduplication, standardization, and schema enforcement typically run as gated checks positioned right at the Bronze-to-Silver boundary, evaluating each record or batch against centrally governed rules while it is still in flight. Records that pass continue into Silver, and records that fail are gated, quarantined, or blocked outright, so bad data never reaches whatever depends on Silver downstream. These are not a separate set of pipeline-specific rules built just for this purpose; they are the same governed, business-defined rules already exposed as gates and re-used by engineers, used to monitor data at rest elsewhere in the estate, and applied at a different point in the data’s path. Deduplication at this stage usually extends into full matching and merging, similar to how master data management handles it elsewhere in an organization: records that represent the same real-world entity, even when they are not exact duplicates, get linked through configurable matching rules and consolidated into a single golden record, with defined logic resolving which value wins when two merged records disagree.

At Gold, the criteria shift from whether the data is clean to whether it can be trusted for production use, and answering that well means combining several signals into a single score across six categories: data quality dimensions like accuracy, completeness, consistency, and validity; observability signals such as anomaly detection and schema drift; lineage and governance transparency; metadata completeness; ownership accountability; and AI-assisted remediation status. A dataset advances to Gold, and becomes discoverable in a data catalog as a governed data product. From there, it serves as the recommended asset behind a business concept like Customer or Transaction, with a named owner and a live Data Trust Index score that shows people and AI agents whether it meets the quality and governance standards their use case requires, and why. Producing that kind of certification, rather than just reporting on data quality after something has already gone wrong, is the job of an end-to-end data quality platform built as part of a broader data trust layer, one designed to help organizations accelerate AI, reduce risk, and modernize data rather than simply flag problems once they have already reached a dashboard.

Running underneath all three layers, data observability catches what rules alone cannot anticipate: a volume anomaly, a distribution that has drifted from its normal range, a schema that changed overnight without anyone flagging it, a job that ran successfully and wrote nothing at all. Rules catch the failure modes someone thought to write a rule for. Observability catches the rest, and the two work as complements rather than substitutes for each other.

The loop closes back at Bronze. Profiling produces the rule conditions that get deployed both as gates at Silver and as ongoing monitoring across the wider estate, and the trust score calculated at Gold is how a team actually verifies whether those gates are doing their job.

Challenges and limitations of medallion architecture

Adopting a three-layer pattern is not free, and the tradeoffs are worth stating plainly rather than glossing over:

  • Storage costs add up. Keeping raw, validated, and aggregated copies of the same underlying data means paying for three versions of it rather than one.
  • Pipeline complexity increases, since each layer transition is another job to build, monitor, and maintain, and that overhead is real even when the pattern is the right choice for the use case.
  • Streaming use cases feel the latency. Waiting for a record to clear Silver-layer validation adds delay that a genuinely real-time use case may not be able to absorb.
  • The pattern can be over-engineered for a low-volume, simple use case that does not need three layers of separation, and building it anyway adds maintenance burden without a matching benefit.
  • Gold can become a bottleneck if promotion from Silver depends on a manual review queue rather than automated, rule-based gating, since Gold then moves only as fast as its slowest reviewer.

Is medallion architecture ELT or ETL?

Short answer: primarily ELT. Raw data loads into Bronze before any transformation happens, and the transformation work runs progressively across Silver and Gold rather than before landing. The pattern is flexible enough to accommodate ETL-style transformation ahead of the load too, depending on the tooling in use, but the medallion pattern itself assumes load first and transform after.

Platform notes: Medallion architecture on Databricks, Snowflake, Azure, and Microsoft Fabric

Medallion architecture itself is platform-agnostic. The pattern does not belong to any one vendor, and the notes below describe how it tends to get implemented on a few major platforms rather than endorsing one over another.

On Databricks, quality checks can integrate directly with Lakeflow and Delta Live Tables pipelines, evaluating data at defined checkpoints before it advances to the next layer, with governed rules translated into SQL and pushed down to run natively on Databricks compute. On Snowflake, a data trust layer like Ataccama can validate data at ingestion and certify datasets before they reach AI agents built on the platform, publishing a trust score those agents can act on directly. Lakehouses hosted in Microsoft Fabric can connect through the platform’s SQL analytics endpoint, extending the same governed rules and business glossary into Fabric without standing up a separate toolchain. And on Azure, Google BigQuery, and other environments, the same underlying rule and governance layer connects through a broader connector footprint spanning more than a hundred out-of-the-box connectors across AWS, Azure, Google Cloud, and Snowflake, alongside on-premises systems and mainframes. That breadth is the point of calling the pattern platform-agnostic in the first place: it’s about the three-layer structure, not with whichever platform happens to be underneath it.

What about a platinum layer?

Some teams extend the standard three layers with a fourth, often called Platinum, sitting above Gold to hold curated semantic models or feature stores built specifically for machine learning. The term is useful to recognize if it comes up, though it is not a standardized part of the pattern the way Bronze, Silver, and Gold are, and most organizations get everything they need from the three-layer structure without adding a fourth.

Data governance and lineage across layers

Every transition between layers is a governance event, whether or not a team treats it that way. Moving data from Bronze to Silver can change its schema, its access controls, and who is accountable for what happens to it next, and none of that should happen implicitly. Lineage tooling that automatically scans metadata across sources, transformation jobs, and destinations, building a visual map of the whole pipeline enriched with quality insights and anomaly alerts at each step, is what makes root-cause analysis and audit trails tractable once a pipeline has more than a handful of layers and jobs. For financial services organizations specifically, that kind of traceability is close to a regulatory requirement rather than a nice-to-have, mapping directly to frameworks like BCBS 239. A data governance program that only checks in once data reaches Gold has already missed the two layers where most of the actual risk gets introduced or resolved.

Frequently asked questions

What are the three layers of medallion architecture?

Bronze, Silver, and Gold. Bronze holds raw data exactly as it arrived from source systems, Silver holds validated, deduplicated, and standardized data, and Gold holds aggregated, business-ready data modeled for direct consumption by BI tools, machine learning models, and AI agents.

What are the downsides of using a medallion architecture?

The main tradeoffs are the storage cost of keeping multiple copies of the same data, added pipeline complexity to build and maintain, latency in streaming contexts where validation introduces delay, the risk of over-engineering a three-layer pattern for a use case that does not need it, and a Gold layer that can become a bottleneck if promotion between layers depends on manual review rather than automated gating.

Is medallion architecture ELT or ETL?

Primarily ELT. Raw data loads first and gets transformed progressively across the layers, though the pattern can accommodate ETL-style transformation before landing depending on the tooling involved.

Who invented medallion architecture?

The pattern originated at Databricks as part of the Delta Lake and Databricks Lakehouse Platform ecosystem, where it is also known as multi-hop architecture.

Does Snowflake use medallion architecture?

Yes. The pattern is platform-agnostic and works across Snowflake, Azure Synapse, Google BigQuery, Microsoft Fabric, and on-premises or hybrid environments alike.

What is the difference between Bronze, Silver, and Gold layers?

Bronze is raw and unmodified, Silver is validated and conformed to a consistent schema, and Gold is aggregated, modeled, and certified as ready for business use, with each layer representing a higher level of trust than the one before it.

What is medallion architecture used for?

Common uses include business intelligence and reporting, machine learning and AI feature pipelines, and regulatory reporting: anywhere an organization needs data to move from raw and unreliable to trusted and analytics-ready in a repeatable way.

Medallion architecture gives a lakehouse its structure. What makes the data within the layers trustworthy is the quality work happening between the layers, including the data quality gates, profiling, and certification, done consistently rather than left to whichever team touches the pipeline last. The earlier that work happens, the cheaper it stays: a defect caught at Bronze is a contained problem, and the same defect caught after it has already reached Gold means tracing a much longer trail back to its source. Shift-left data quality: Why downstream checks are too late for data products and AI goes deeper into what that shift actually requires, from source systems through the pipeline to the lakehouse itself. 

Author

Joellen Koester

JoEllen is the Director of Content Strategy at Ataccama and has worked in the AI and data spaces since 2015. She holds bachelor's degrees in English and Philosophy from Seattle University, a master's degree in Transatlanic Studies from Charles University, and was awarded a Fulbright Scholarship to teach English in the Czech Republic.

Published at 05.10.2026

Do you like this content?
Share it with others.