Shanku Niyogi, Vice President of Product Management at Databricks, has introduced a provocative reinterpretation of a foundational industry acronym, labeling the traditional Change Data Capture (CDC) process as "continuous data corruption." Speaking during the week of the Databricks Data and AI Summit in San Francisco, Niyogi argued that the streaming pipeline techniques used for decades to shuttle operational data into analytical warehouses are no longer fit for purpose in an era dominated by artificial intelligence and high-velocity application development. The critique serves as the philosophical cornerstone for Databricks’ latest technological pivot: the launch of an "agentic data foundation" designed to collapse the wall between transactional and analytical systems.
The traditional CDC model operates by piping copies of every change in a transactional database—the systems handling real-time orders, payments, and inventory—into an analytics warehouse. This separation was historically necessary because the two types of databases were optimized for fundamentally incompatible tasks. However, Niyogi contends that this workaround has become a liability for modern data engineers. According to Niyogi, CDC pipelines are notoriously slow, expensive, and prone to failure when schemas change or network latency intervenes. By rebranding the process as "continuous data corruption," Databricks is highlighting the inherent "pain" felt by data teams who must maintain brittle bridges between disparate data silos.
The Rise of the Agentic Data Foundation
To address these systemic inefficiencies, Databricks has unveiled a two-pronged strategy involving Lake Transactional/Analytical Processing (LTAP) and Lakehouse//RT. These innovations represent a shift toward a unified architecture where transactional and analytical data coexist within a single framework, governed by a unified system. This "agentic data foundation" is specifically tailored to support the next generation of AI-driven applications, which require the ability to reason over and act upon data in near real-time without the delays inherent in traditional Extract, Transform, Load (ETL) processes.
The urgency of this shift is underscored by the explosive growth in software development. Niyogi noted that the volume of code being written globally has increased 50-fold over the past year. Projections suggest that within the next 12 months, the world will produce more code than in the entirety of human history. This surge is driven largely by AI agents that not only write code but also power the applications themselves. These applications demand a data environment that can handle high-frequency writes and complex analytical queries simultaneously—a feat that traditional CDC-dependent architectures struggle to achieve at scale.
LTAP and the Evolution of Lakebase
At the heart of this new architecture is Lakebase, a product Databricks launched a year ago following its acquisition of the Postgres-compatible database company Neon in 2025. Lakebase utilizes a design philosophy that separates compute from storage, a standard practice in modern data warehousing that is now being applied to transactional workloads. While operational databases traditionally require local disk storage to maintain sub-millisecond latency, Lakebase leverages cloud object storage—such as Amazon S3, Microsoft Azure Blob, and Google Cloud Storage—to provide infinite scalability at a lower cost.
The technical breakthrough of LTAP lies in how it handles data ingestion. When a Postgres database within Lakebase writes a row—the standard format for operational tasks—the system simultaneously converts that data into a columnar format, such as Delta Lake or Apache Iceberg. These open table formats are then stored directly on the data lake. This dual-format writing process ensures that data becomes available for analytical querying within minutes. By eliminating the need for a physical reconciliation pipeline, Databricks ensures that the operational and analytical copies are effectively the same data, governed by a single security and compliance model.
Niyogi emphasized the developer experience within this ecosystem, particularly the introduction of database branching. Similar to how developers use Git to manage code, Lakebase allows for the cloning of a database to run experiments or test AI agent transactions in a safe environment before merging them with production state. This capability is critical for "agentic" workloads where AI models must interact with production data without risking system integrity.
Performance Benchmarks and Lakehouse//RT
The second major component of this announcement is Lakehouse//RT, a real-time query layer built on a new compute engine codenamed "Reyden." This engine is designed to overcome what Niyogi calls the "one-second barrier"—the latency floor that most data warehouses hit when trying to serve application-level workloads. Traditionally, developers have bypassed this limitation by copying data from the warehouse into specialized caching layers like Redis or purpose-built serving databases.
Lakehouse//RT aims to make such workarounds obsolete. Databricks claims that the Reyden engine, which utilizes a fully asynchronous execution model, can provide a 10-millisecond floor for direct queries against the lakehouse. Furthermore, the platform reportedly supports tens of thousands of concurrent users, allowing entire organizations to run high-performance applications directly against the warehouse without data duplication.
Early performance data released by Databricks suggests that Lakehouse//RT is up to 16 times faster than existing specialized real-time serving stacks. While the specific competitors in these benchmarks were not named, early adopters have reported significant gains. Chris Kopek, Head of Data Platforms at Cisco, stated that threat-lookup queries are running five times faster than their previous configuration. Similarly, Kayvon Raphael of Magnite reported sub-200-millisecond performance on core dashboard queries, even at high volumes of hundreds of queries per second.
Chronology of Innovation and Market Adoption
The transition toward a unified agentic foundation is the latest step in a multi-year roadmap for Databricks.
- 2020: Databricks introduces the "Lakehouse" concept, merging data lakes and data warehouses.
- 2024 (Present): The launch of Lakehouse//RT and the formalization of LTAP at the Data and AI Summit.
- 2025: The acquisition of Neon (as per company roadmap) to bolster transactional capabilities via Lakebase.
- Future Outlook: Databricks anticipates a near-total shift in how enterprise applications are architected, moving away from fragmented stacks toward unified, open-format foundations.
The market is already showing signs of rapid adoption. Lakebase currently serves over 3,500 customers, with Databricks reporting 12 million database launches per day across its platform. High-profile clients including Block, Zillow, and Superhuman are utilizing the platform to streamline their data operations. Ensemble, a healthcare revenue cycle management firm, is using the technology to manage over two petabytes of harmonized data to drive AI-led revenue recovery in hospital environments.
Implications for the Enterprise CIO
For Chief Information Officers (CIOs), the message from Databricks is clear: the traditional bifurcated data stack is an obstacle to AI maturity. Niyogi argued that CIOs can no longer afford to "stitch together" systems using processes built decades ago. As organizations move from pilot AI programs to hundreds of production applications, the operational overhead of maintaining thousands of individual CDC pipelines becomes unsustainable.
The strategic shift toward LTAP and Lakehouse//RT is framed as a simplification of the developer’s "mental model." In this new paradigm, Lakebase serves as the destination for writes—fully compatible with existing Postgres tools and drivers—while Lakehouse//RT serves as the high-speed engine for analytical reads. This eliminates the need for ETL, reverse ETL, and complex caching layers, theoretically reducing both cloud spend and engineering headcount requirements.
Fact-Based Analysis and Industry Context
While Databricks’ performance claims are impressive, industry analysts note that the success of this architecture depends on the continued adoption of open standards. Databricks has leaned heavily into Apache Iceberg and Delta Lake, and recently donated its OpenSharing protocol to the Linux Foundation to foster a more interoperable ecosystem. However, the competitive landscape remains fierce. Specialized databases and rival warehouse providers are also racing to reduce latency and bridge the transactional gap.
The claim of a 16-fold performance increase will likely face scrutiny as independent benchmarks emerge. Historically, "all-in-one" solutions have struggled to match the extreme performance of specialized tools like Redis for caching or ClickHouse for real-time analytics. However, Databricks’ move to integrate these capabilities directly into the lakehouse suggests a bet that "good enough" real-time performance, combined with superior governance and lower complexity, will win over enterprise architects.
Furthermore, the "agentic" focus reflects a broader trend in the industry. As S&P Global Market Intelligence reports that approximately 31% of organizations already have at least one AI agent in production, the demand for "agent memory"—the ability for an AI to quickly write to and read from a persistent state—is becoming a primary driver for database selection. By positioning Lakebase and Lakehouse//RT as the ideal "memory" for these agents, Databricks is attempting to capture the foundational layer of the burgeoning AI economy.
Ultimately, the move to replace "continuous data corruption" with a unified transactional-analytical engine represents a significant gamble on the future of data engineering. If Databricks can deliver on its promise of sub-second lakehouse queries and automated row-to-columnar conversion, it may fundamentally alter the blueprint for enterprise data architecture in the age of artificial intelligence.
