A working catalog on agents, semantics, and context.
Essays and guides on data engineering agents, semantic layers, MCP, and text-to-SQL.
Latest
The most recently published and updated posts.
Datus: The Cursor for Data Engineering
"Cursor for data engineering" means an agent that runs your data system — warehouses, metrics, and reference SQL — with evolvable context, not autocomplete.
Open Semantic Interchange (OSI): What the New Standard Means for Data Engineering and AI Agents
A complete guide to Open Semantic Interchange (OSI) — now Apache Ossie (incubating): what it standardizes, who's behind it, and why portable semantics matter.
Semantic Layer Tools in 2026: Complete List + OSI (Apache Ossie) Status
Every semantic layer tool in 2026 — dbt MetricFlow, Cube, AtScale, Snowflake, LookML and more — with each one's current OSI (Apache Ossie) support status.
What Is Change Data Capture (CDC)? Methods & Use Cases
Change data capture (CDC) definition, the three CDC methods — log-based, trigger, query — plus real-time use cases, pitfalls, and CDC to the lakehouse.
What Is a Data Contract? Definition, Schema Enforcement & Examples
Data contract definition: a machine-checked producer–consumer agreement on schema, semantics, quality, and SLAs. Tools, enforcement, and the agent angle.
What Is Medallion Architecture? Bronze, Silver & Gold Layers
Medallion architecture definition: Bronze, Silver, and Gold lakehouse layers, what belongs in each, the anti-patterns, and which layer an AI agent should query.
What is Datus
Start here — the problem, the product, and the thesis behind it.
Datus: The Cursor for Data Engineering
"Cursor for data engineering" means an agent that runs your data system — warehouses, metrics, and reference SQL — with evolvable context, not autocomplete.
Meet the General Chat Agent: Your Data Co-Pilot That Actually Thinks
The General Chat Agent goes beyond SQL generation to support exploration, investigation, and knowledge-building.
SQL agents are broken without context. Meet Datus.
Learn why SQL agents fail without governed context and how Datus uses contextual engineering and subagents for reliable workflows.
From Human-First Data Systems to the Agentic Data Stack
Learn why the Agentic Data Stack goes beyond AI SQL tools by combining context, semantics, workflows, and governed execution.
Welcome to Datus Blog
Introducing the Datus blog with insights on AI-native data engineering, context engineering, and reliable data workflows.
Data Engineering Agent
The category, the comparisons, and how to build with one — our core cluster.
How Datus Turns AI-Generated SQL into Trusted Data
Why reliable AI data engineering needs knowledge, planning, review, controlled execution, and reconciliation, not just SQL generation.
What Is a Data Engineering Agent? Definition, Examples & a 2026 Comparison
Four products now ship as a data engineering agent — but they are not the same thing. Working definition, side-by-side comparison, and where persistent context separates agents from chat windows.
What Is a Data Engineering Agent? A Practical Guide with Datus
Learn what a data engineering agent is, why context matters, and how Datus turns AI into reliable, production-ready data workflows.
Contextual Data Engineering: Why Every Data Engineering Agent Needs Evolvable Context
Contextual data engineering explained: schemas, semantics, and feedback loops for durable data agents.
Best Data Engineering Agents in 2026: An Honest Comparison
Best data engineering agents in 2026 compared by stack fit, context, openness, and enterprise readiness.
Open Source Data Engineering Agents: Why They Exist, When to Use One, and What Your Options Are
Open-source data engineering agents compared: Datus, Wren AI, Altimate, and when self-hosting is worth it.
How to Build Your First Data Engineering Agent in 15 Minutes
Build a first data engineering agent with Datus: install, ask questions, generate context, and create a subagent.
Data Engineering Agent vs. Claude Code: When to Use Which
Data engineering agent vs Claude Code: when persistent data context matters and when a coding agent is enough.
Data Engineering Agent vs. SQL Copilot: What's the Real Difference?
Data engineering agent vs SQL copilot: persistence, feedback, team context, and when each tool fits.
One-Person Data Team: How a Data Engineering Agent Multiplies Your Output
How a one-person data team uses a data engineering agent to reduce SQL translation work and ship self-service analytics.
How a Context Engine Makes Data Engineering Agents More Accurate
How a context engine improves data engineering agent accuracy with schemas, validated SQL, and feedback loops.
MCP and Data Engineering: The Protocol That Connects Your Entire Stack
MCP for data engineering: how agents connect to databases, orchestrators, quality tools, and context services.
What an Enterprise Data Engineering Agent Actually Needs
Enterprise data engineering agent requirements: shared context, RBAC, auditability, reliability, and governance.
Subagents: How to Ship Domain-Specific Data Agents Without Training a Model
Subagents explained: domain-specific data agents built from scoped context, feedback, and governed access.
Best Data Engineering Agents in 2026: An Honest Comparison
Best data engineering agents in 2026 compared by stack fit, context, openness, and enterprise readiness.
AI-Native Data Platforms: Why the Next Generation Needs Data Engineering Agents, Not Just Copilots
What defines an AI-native data platform, how it differs from platforms with bolted-on AI features, and why data engineering agents are the missing infrastructure layer.
Platform-Native Data Engineering Agents Compared: Cortex Code, Genie Code, and BigQuery DE Agent
A detailed comparison of Snowflake Cortex Code, Databricks Genie Code, and Google BigQuery Data Engineering Agent — and the case for open, cross-stack alternatives.
Semantic Layer
What a semantic layer is, and how it differs from a metric layer, model, ontology, or catalog.
What Is a Semantic Layer? Definition, Examples & How It Differs From a Metric Layer
Semantic layer defined: the business translation layer between raw tables and analysts, what it includes (metrics, dimensions, entities), how it differs from metric layers and catalogs, and why static models break under AI agents.
What Is a Metric Layer? Definition, Examples & How It Differs From a Semantic Layer
Metric layer definition, MetricFlow examples, semantic layer vs metric layer differences, and why AI agents need standardized metrics.
What Is a Semantic Model? Definition, Examples & How It Differs From a Semantic View
Semantic model definition, key components, how it fits into a semantic layer, and how it differs from warehouse-native semantic views.
Semantic Layer vs Ontology: What's the Difference and Why It Matters for AI Agents
How semantic layers and ontologies relate, where they diverge, and why understanding both matters for building AI agents that can trust data.
Open Semantic Interchange (OSI): What the New Standard Means for Data Engineering and AI Agents
A complete guide to Open Semantic Interchange (OSI) — now Apache Ossie (incubating): what it standardizes, who's behind it, and why portable semantics matter.
OSI vs MetricFlow: Semantic Standard vs Execution Engine
OSI vs MetricFlow: Open Semantic Interchange is the portable semantic standard; MetricFlow is dbt's execution engine—how they differ and when to use each.
dbt Semantic Layer & MetricFlow: A Complete Guide for Data Engineers
How dbt Semantic Layer and MetricFlow work, what they mean for data engineering teams, and how they fit with data engineering agents.
Cube.dev: From Semantic Layer Pioneer to Agentic Analytics Platform
How Cube.dev evolved from an open-source semantic layer to the D3 Agentic Analytics platform, and what its trajectory means for data engineering.
GoodData: How a 17-Year BI Company Became an AI-Native Analytics Platform
GoodData's evolution from cloud BI startup to GoodData.AI — what it reveals about the industry shift toward AI-native analytics and the role of the semantic layer.
Semantic Layer Tools in 2026: Complete List + OSI (Apache Ossie) Status
Every semantic layer tool in 2026 — dbt MetricFlow, Cube, AtScale, Snowflake, LookML and more — with each one's current OSI (Apache Ossie) support status.
Glossary
Core data engineering terms — defined, with how they connect to agents and context.
What Is Text-to-SQL? Definition, How It Works & Why Context Matters
Text-to-SQL definition, NL2SQL pipeline stages, accuracy limits, and how data engineering agents improve with persistent context.
What Is Schema Linking? Definition, Challenges & How Agents Map NL to Columns
Schema linking definition for text-to-SQL, common failure modes, and how dual-dimension context improves column resolution.
What Is RAG for Data Engineering? Retrieval, Context & Agent Accuracy
RAG definition for data engineering: retrieving schema, metrics, and SQL history to ground NL2SQL and data engineering agents.
What Is a Data Catalog? Definition, Tools & How It Differs From Agent Context
Data catalog definition, popular tools, and why data engineering agents need context engines beyond discovery metadata.
What Is Data Mesh? Definition, Principles & How Domain Agents Map to It
Data mesh definition, four principles, comparison to data fabric, and how subject trees and subagents align with domain ownership.
What Is a Data Agent? How It Differs From a Data Engineering Agent
Data agent definition, types, capabilities, and how a data engineering agent fits as the specialized subclass that builds and evolves data context.
What Is a Lakehouse? Definition, Architecture & Open Table Formats Explained
Lakehouse definition, how it differs from data lakes and warehouses, open table formats (Iceberg, Delta, Hudi), and why AI agents need lakehouse-aware context.
What Is a Lakehouse Catalog? Hive, Glue, Unity, Polaris & Horizon
A lakehouse catalog tracks table metadata for query engines. Compare Hive Metastore, AWS Glue, Unity Catalog, Apache Polaris and Snowflake Horizon.
What Is a Data Warehouse? Definition, Architecture & How It Differs From a Data Lake
Data warehouse definition, architecture (ETL, dimensional models, columnar MPP), how it differs from a data lake and lakehouse, and what AI agents need to query one.
What Is a Data Lake? Definition, Architecture & Data Lake vs Data Warehouse
Data lake definition, schema-on-read architecture, zones and file formats, the data swamp problem, data lake vs data warehouse, and what AI agents need to query one.
What Is a Data Contract? Definition, Schema Enforcement & Examples
Data contract definition: a machine-checked producer–consumer agreement on schema, semantics, quality, and SLAs. Tools, enforcement, and the agent angle.
What Is Medallion Architecture? Bronze, Silver & Gold Layers
Medallion architecture definition: Bronze, Silver, and Gold lakehouse layers, what belongs in each, the anti-patterns, and which layer an AI agent should query.
What Is Change Data Capture (CDC)? Methods & Use Cases
Change data capture (CDC) definition, the three CDC methods — log-based, trigger, query — plus real-time use cases, pitfalls, and CDC to the lakehouse.
Why agents, not copilots
The shift from assistive AI to autonomous data workflows.
Why Data Engineering Needs Agents, Not Just Copilots
Learn why data engineering needs agents, not just copilots, and how agentic workflows improve execution, reliability, and control.
Agentic Data Engineering vs Traditional Data Engineering
Compare agentic and traditional data engineering across workflows, tooling, team structure, and reliability in production.
What Autonomous Data Engineering Actually Looks Like in Practice
See what autonomous data engineering looks like in practice with structured context, bounded agent actions, human review, and workflow automation.
Architecture and pipelines
How agentic data systems are built and run.
Datus Storage Layer: A Foundation Built for Every Environment
Pluggable storage adapters that separate relational and vector storage from the agent core for enterprise flexibility.
Data Engineering Agent Architecture for Production
A practical architecture blueprint for data engineering agents, with Datus patterns for context, subagents, and governed execution.
AI Data Pipeline Automation: Use Cases, Architecture, and Tradeoffs
Learn where AI data pipeline automation works, how to design the architecture, and which tradeoffs matter before scaling it.
Agentic ETL: What Changes Beyond Traditional ETL
See how agentic ETL adds context-aware planning, validation, workflow-state reasoning, and human review beyond traditional ETL.
Why context is everything
Why durable, structured context is what makes agents reliable.
Why AI Agents Need Semantic Context to Work Reliably
Learn why semantic context helps AI agents reason reliably across data systems by grounding definitions, relationships, and constraints.
How Structured Context Improves AI Agent Output
Learn how structured context improves AI agent output by grounding reasoning in metrics, semantics, and workflow state.
Semantic Modeling for Agentic Analytics Workflows
Learn how semantic modeling improves agentic analytics by grounding AI agents in shared metric definitions and governed business context.
Why Reliable Data Agents Need More Than Good Prompts
Learn why reliable data agents need structured context, workflow state, metrics, semantic models, and guardrails beyond good prompts.
Tooling and integrations
How Datus connects to the rest of your stack.
How MCP Changes Data Workflow Automation
Learn how MCP improves data workflow automation with structured tool access, safer execution paths, and tighter system integration.
Using MCP Extensions in Data Engineering Workflows
Learn how MCP extensions give data engineering agents controlled tool access, safer execution paths, and more reliable automation.
Beyond SQL: How Datus Integrates With Your Entire Data Toolchain
Learn how MCP and Skills connect Datus to your data catalog, metric layer, scripts, and quality workflows across the full data toolchain.
In practice
Use cases and operating models from real teams.
7 High-Impact Data Engineering Agent Use Cases (Powered by Datus)
Explore practical data engineering agent use cases and how Datus helps teams improve speed, quality, and governance.
The Operating Model of an Agentic Data Team
Learn how an agentic data team operates with clear roles, review loops, guardrails, and human-in-the-loop control across planning, execution, and governance.
Make Data Agents Usable: Ask, Explore, and Control with Confidence
See how Ask User, session management, Explore, and action display make Datus data agents easier to trust, control, and use every day.
Releases
What's new in Datus.
More essays
What Is Apache Hudi? Upserts, Copy-on-Write vs Merge-on-Read & CDC
Apache Hudi definition, how copy-on-write and merge-on-read tables, record-level upserts and incremental queries work, Hudi vs Iceberg vs Delta, and agent context.
What Is Apache Iceberg? Table Format, Features & Iceberg vs Delta Lake
Apache Iceberg definition, how its hidden partitioning, snapshots and schema evolution work, Iceberg vs Delta Lake vs Hudi, and why AI agents need table-aware context.