Data Platform Reference Architecture
A comprehensive reference architecture for building a modern, cloud-native data platform — covering ingestion, storage, processing, consumption, governance, security, and DataOps with product recommendations across all major cloud vendors.
Data Platform Component Model
Hover over components to highlight, and click any item to see capability summaries and cloud vendor services.
Overview
This reference architecture defines a comprehensive, generic Data Platform designed to scale seamlessly and break down organizational data silos. Structured from the foundational infrastructure up to the end consumers, it outlines the essential capabilities required to ingest, store, process, and serve data, wrapped in robust governance, security, and operational pipelines.
Component Model Layers
1. Data Sources Layer
This foundational layer represents the diverse origins of data that exist before entering the platform. These sources generate the raw information that feeds the enterprise.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Operational Databases | Core transactional relational and document engines backing the primary operational products of the enterprise. | Amazon RDS / Aurora (AWS); Azure SQL / PostgreSQL (Azure); Cloud SQL / Cloud Spanner (GCP); PostgreSQL, MySQL, MongoDB (Open Source) |
| Applications | Business enterprise software suites, cloud SaaS applications, CRM, ERP, and general line-of-business software platforms. | Salesforce Integration (AWS); SAP on Azure (Azure); Workday on GCP (GCP); Odoo, ERPNext (Open Source) |
| Streams | Unbounded telemetry feeds, clickstream trails, and internet-connected devices transmitting active continuous sensor measurements. | AWS IoT Core (AWS); Azure IoT Hub (Azure); Google Cloud IoT (GCP); Eclipse Mosquitto, EMQX (Open Source) |
| Structured Data | Highly structured and tabular data formatted files sent by internal teams or external vendors (such as monthly CSV summaries). | Amazon S3 Transfer (AWS); Blob Storage Inbound (Azure); GCS Bucket Uploads (GCP); SFTP, Samba (Open Source) |
| Unstructured Data | Binary, heavy, and media files containing rich operational data lacking pre-defined relational structures (e.g. PDFs, images). | Amazon S3 Object Store (AWS); Azure Blob Storage (Azure); Google Cloud Storage (GCP); MinIO Object Storage (Open Source) |
2. Infrastructure Layer
The underlying physical or virtualized hardware and cloud services that host the data platform. This layer provides the foundational compute, storage, and networking capabilities on which all other layers depend.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Cloud Infrastructure | Primary scalable virtual machines, bare metal components, and block storage platforms serving as the core host of the platform. | Amazon EC2 & EBS (AWS); Azure VMs & Managed Disk (Azure); Compute Engine & Persistent Disk (GCP); OpenStack (Open Source) |
| Container Orchestration | Schedules, scales, and manages containerized applications and transformation tasks consistently across hybrid environments. | Amazon EKS (AWS); Azure Kubernetes Service (AKS) (Azure); Google Kubernetes Engine (GKE) (GCP); Kubernetes, Nomad (Open Source) |
| Private Networking | Ensures isolation of cloud workloads using virtual firewalls, private end-points, and secure virtual private tunnels to block exfiltration. | AWS VPC & PrivateLink (AWS); Azure Virtual Network & Private Link (Azure); Google Cloud VPC & PSC (GCP); Cilium, Calico (Open Source) |
| On-Premises / Hybrid | Secure on-premises physical data centres linked via high-speed dedicated tunnels for low-latency or strict data residency systems. | AWS Outposts / Direct Connect (AWS); Azure Stack / ExpressRoute (Azure); Google Anthos / Interconnect (GCP); VMware, Nutanix (Open Source) |
3. Data Ingestion Layer
This layer extracts data from diverse source systems, acting as the primary consumer of the Enterprise Integration Platform for event-driven and API-based feeds. Note on Data Contracts: Across all ingestion methods, this layer acts as the enforcement point for Data Contracts—ensuring incoming data adheres to agreed-upon schemas and SLAs before it enters the platform.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Database Ingestion | Safely reads transactional logs in real-time (Change Data Capture) or executes batches without impacting database performance. | AWS Database Migration Service (AWS); Azure DB Migration (Azure); GCP Database Migration (GCP); Debezium, Fivetran, Airbyte (Open Source) |
| Stream & Messaging Ingestion | Directly ingests, aggregates, and pipes high-velocity messaging topics and streaming data sources straight into storage. | Amazon Kinesis Data Streams (AWS); Azure Event Hubs (Azure); Google Pub/Sub (GCP); Apache Kafka, Logstash (Open Source) |
| API Ingestion | Connects securely to third-party SaaS tools and web endpoints to pull or push data assets via REST/GraphQL or webhooks. | Amazon AppFlow (AWS); Azure Logic Apps (Azure); Google Cloud Data Fusion (GCP); Meltano, Singer (Open Source) |
| File Ingestion | Managed secure data transfer services to safely receive, move, and store external file exports, flat archives, and XMLs. | AWS Transfer Family (AWS); Azure Data Factory (Azure); GCP Storage Transfer (GCP); rclone, Apache NiFi (Open Source) |
4. Data Storage Layer
This layer provides scalable, decoupled storage for all data types. It supports everything from raw object storage to highly structured analytical models and specialized operational formats.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Data Lake | Highly cost-effective object storage holding all historical records. Modern architectures overlay open table formats (Iceberg, Delta) for ACID reliability. | Amazon S3 + Apache Iceberg (AWS); ADLS Gen2 + Delta Lake (Azure); Google Cloud Storage + Iceberg (GCP); MinIO + Apache Iceberg (Open Source) |
| Enterprise Data Warehouse (EDW) | Structured analytical relational storage optimized for complex reporting queries, running highly performant mass analytical schemas. | Amazon Redshift (AWS); Azure Synapse / Fabric (Azure); Google BigQuery (GCP); Snowflake, ClickHouse (Open Source) |
| Streaming Broker | Distributed, immutable commit log with configurable retention (Kafka-style) that durably stores high-velocity event streams, enabling replay, decoupled consumers, and multi-subscriber fan-out. | Amazon MSK (AWS); Azure Event Hubs (Azure); Google Pub/Sub (GCP); Apache Kafka, Redpanda (Open Source) |
| Purpose-Built Databases | Specialized non-relational database services optimized for document, graph, key-value, or low-latency operational data needs. | Amazon Neptune / DocumentDB (AWS); Azure Cosmos DB (Azure); Google Firestore / Spanner (GCP); MongoDB, Neo4j, PostgreSQL (Open Source) |
| Vector Store | Specialized databases designed to store, index, and query high-dimensional vector embeddings—crucial for modern semantic search, RAG, and generative AI pipelines. | Amazon OpenSearch / RDS pgvector (AWS); Azure Cosmos DB (Vector Search) / AI Search (Azure); Vertex AI Vector Search / Cloud SQL pgvector (GCP); pgvector, Qdrant, Milvus, Chroma (Open Source) |
5. Data Processing Layer
This layer encompasses the compute engines responsible for transforming, cleaning, enriching, and modeling the data. It also includes the heavy computational lifting required for artificial intelligence.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Distributed Batch Processing | Scalable parallel compute clusters executing scheduled transformations on vast historical datasets. | Amazon EMR, AWS Glue (AWS); Azure Databricks, Synapse Spark (Azure); Google Dataproc (GCP); Apache Spark, dbt (Open Source) |
| Stream Processing | Real-time stream computations transforming unbounded data flows instantly for alerting and rapid-response dashboarding. | Kinesis Data Analytics (Flink) (AWS); Azure Stream Analytics (Azure); Google Dataflow (GCP); Apache Flink, Spark Streaming (Open Source) |
| ML Model Training | High-performance GPU-backed compute used to develop, train, evaluate, and version machine learning and deep learning models — distinct from inference, which is served via the Consumption layer. | Amazon SageMaker (AWS); Azure Machine Learning (Azure); Google Vertex AI (GCP); Kubeflow, PyTorch (Open Source) |
6. Data Consumption Layer
This layer exposes the processed data and trained models to the end consumers. Note on Data Contracts: Interfaces exposed here (like APIs and outbound feeds) are governed by outbound Data Contracts, guaranteeing reliability for downstream consumers.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Data Queries | Ad-hoc, federated SQL query engines that allow interactive querying directly on the data lake without moving it to a warehouse. | Amazon Athena (AWS); Synapse Serverless (Azure); BigQuery (External) (GCP); Trino, Presto (Open Source) |
| Data Visualization | Business Intelligence tools used to construct dynamic, interactive dashboards and scheduled, paginated reports. | Amazon QuickSight (AWS); Microsoft Power BI (Azure); Google Looker / Looker Studio (GCP); Apache Superset, Metabase (Open Source) |
| Model Serving & Advanced Analytics | Hosted, highly scalable inference endpoints serving low-latency predictions and generative model results to client systems, alongside complex statistical analysis and exploratory data science workloads. | SageMaker Endpoints (AWS); Azure ML Endpoints (Azure); Vertex AI Endpoints (GCP); BentoML, KServe (Open Source) |
| Data Activation / Reverse ETL | Pushes curated analytical profiles (e.g. customer churn risk scores) directly back into operational business applications to trigger automated action. | Amazon AppFlow (AWS); Azure Data Factory (Azure); BigQuery Data Transfer (GCP); Hightouch, Census (Open Source) |
| Operational Data Stores (ODS) | Read-optimized serving stores that cache curated, pre-aggregated data products for low-latency retrieval by customer-facing applications — bridging the analytical platform and operational systems. | Amazon DynamoDB / ElastiCache (AWS); Azure Cache for Redis (Azure); Google Cloud Bigtable (GCP); Apache Cassandra, Redis (Open Source) |
| Data APIs | Exposes secure GraphQL or REST APIs on top of data warehouses or operational caches to grant programmatic access to application developers. | Amazon API Gateway (AWS); Azure API Management (Azure); Google Apigee (GCP); Hasura, PostgREST (Open Source) |
7. Data Consumers Layer
The topmost layer, representing the end-users and systems that derive business value from the data platform.
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Business Users | Analysts, executives, and operational business staff relying on visual dashboards, metrics, and BI reports for daily strategic decisions. | Amazon QuickSight (AWS); Microsoft Power BI (Azure); Google Looker (GCP); Apache Superset (Open Source) |
| Data Scientists | Power users building machine learning models, conducting advanced predictive analytics, and running exploratory data experiments. | SageMaker Notebooks (AWS); Azure ML Studio (Azure); Vertex AI Workbench (GCP); JupyterHub, MLflow (Open Source) |
| Applications | Internal or custom line-of-business applications consuming operational or analytical data products to drive workflows. | AWS AppSync (AWS); Azure App Service (Azure); Google App Engine (GCP); Next.js / Node.js (Open Source) |
| External 3rd Parties | Third-party clients, external vendors, or industry partners consuming platform data safely via secure outbound B2B feeds or API gateways. | AWS Data Exchange (AWS); Azure Data Share (Azure); Analytics Hub (GCP); Delta Sharing (Open Source) |
| Automated Systems | Autonomous systems, enterprise robotic process automation (RPA), and automated algorithms consuming insights with no human in the loop. | AWS Step Functions (AWS); Azure Logic Apps (Azure); GCP Workflows (GCP); Node-RED, temporal.io (Open Source) |
| Technical Operations | Data engineers, platform support teams, and SRE professionals monitoring pipeline stability, data quality alerts, and costs. | AWS CloudWatch (AWS); Azure Monitor (Azure); Google Cloud Operations (GCP); Grafana, Prometheus (Open Source) |
Cross-Cutting Layers
These layers span horizontally across the platform, ensuring data remains secure, compliant, understandable, and operations run smoothly.
Data Governance Layer
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Data Discovery & Metadata | A searchable database of data schemas, terminology definitions, and ratings allowing users to find data fast. | AWS Glue Data Catalog (AWS); Microsoft Purview (Azure); Google Dataplex Catalog (GCP); Amundsen, DataHub, Apache Atlas (Open Source) |
| Data Quality | Continuously asserts quality parameters (freshness, completeness, shape) on active pipelines and alerts on deviation. | AWS Glue Data Quality (AWS); Purview Data Quality (Azure); Google Dataplex DQ (GCP); Great Expectations, Soda, Deequ (Open Source) |
| Data Lineage | Draws visual execution pathways of data from original source, through processing transformations, to visualization dashboards. | AWS Glue Lineage (AWS); Microsoft Purview Lineage (Azure); Google Dataplex Lineage (GCP); OpenLineage, DataHub (Open Source) |
| Master & Reference Data Management | Reconciles and creates a single golden master record for core business entities (like 'Customer' or 'Product'). | Systems Integration (AWS); Profisee on Azure (Azure); Reltio on GCP (GCP); TIBCO EBX, Talend MDM (Open Source) |
Data Security Layer
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Access Control | Enforces row-level, column-level, or cell-level access constraints on queries depending on corporate role clearance. | AWS Lake Formation, IAM (AWS); Microsoft Purview Access (Azure); BigQuery Policy Tags, IAM (GCP); Apache Ranger, Immuta (Open Source) |
| Data Masking & Encryption | Obfuscates sensitive parameters like credit card numbers dynamically and encrypts data both in transit and at rest. | AWS KMS, Redshift Masking (AWS); Azure Key Vault, SQL Masking (Azure); Google Cloud KMS, DLP API (GCP); HashiCorp Vault, Apache Ranger (Open Source) |
| Secrets Management | Central repository holding, rotating, and tracking programmatic access credentials and server connection strings safely. | AWS Secrets Manager (AWS); Azure Key Vault Secrets (Azure); Google Secret Manager (GCP); HashiCorp Vault (Open Source) |
| Data Classification & Sensitive Data Discovery | Scans new data objects to identify and tag PII (Personally Identifiable Information), health tags, or banking numbers. | Amazon Macie (AWS); Microsoft Purview Information Protection (Azure); Google Cloud Sensitive Data Protection (GCP); Apache Ranger Classification (Open Source) |
| Audit Logging | Tamper-evident system capture documenting every single query, configuration alteration, and file download action. | AWS CloudTrail (AWS); Azure Monitor Logs (Azure); Google Cloud Audit Logs (GCP); Elasticsearch, OpenSearch (Open Source) |
| Identity & Access Management (IAM) | Federates user profiles, handles identity authentication, and issues safe machine credential tokens to control who and what can access platform resources. | AWS IAM & Identity Center (AWS); Microsoft Entra ID (Active Directory) (Azure); Google Cloud IAM (GCP); Keycloak, Okta (Open Source) |
Supporting Services (DataOps) Layer
| COMPONENT | DESCRIPTION | EXAMPLES |
|---|---|---|
| Data Orchestration & Workflow Management | Pipes execution dependency commands across tasks, ensuring batches complete sequentially and retry on failure. | Amazon MWAA (Airflow) (AWS); Azure Data Factory (Azure); Google Cloud Composer (GCP); Apache Airflow, Prefect, Dagster (Open Source) |
| Monitoring & Observability (incl. FinOps) | Tracks resource workloads, queries execution cost metrics, system crashes, and platform performance latency in real-time. | Amazon CloudWatch & Cost Explorer (AWS); Azure Monitor & Cost Management (Azure); GCP Operations & FinOps (GCP); Grafana, Prometheus, Datadog (Open Source) |
| CI/CD & Source Control | Automates testing code integrations, model versioning, and environment container migrations directly from Git. | AWS CodePipeline (AWS); Azure Pipelines (Azure); Google Cloud Build (GCP); GitHub Actions, GitLab CI (Open Source) |
| Infrastructure as Code (IaC) | Represents data warehouses and infrastructure configurations as code structures, allowing rapid environment duplication. | AWS CloudFormation (AWS); Azure Bicep (Azure); GCP Deployment Manager (GCP); Terraform, OpenTofu, Pulumi (Open Source) |
| Data Modeling Tools | Visual and programmatic systems to map relational schemas, relational keys, warehouse stars, and data vault architectures. | AWS Glue Schema Registry (AWS); Synapse Database Designer (Azure); BigQuery Schema Editor (GCP); dbdiagram.io, SqlDBM, Erwin (Open Source) |