The modern AI search landscape demands infrastructure capable of processing high-volume, low-latency database queries at scale, prompting artificial intelligence startup Perplexity to bypass managed cloud databases in favor of a custom solution. Facing prohibitive costs and latency constraints with Amazon Web Services’ DynamoDB, Perplexity engineers designed and deployed CobbleDB, a custom 40,000-line Rust key-value store. Developed by a two-person engineering team in just two months with substantial assistance from persistent coding agents, CobbleDB now handles a significant share of Perplexity’s live production search traffic. The homegrown database has achieved a median batch-read latency of 5.6 milliseconds—a dramatic improvement over the 31.4 milliseconds recorded under DynamoDB—while reducing projected operational costs by at least 20 percent.
The project highlights a broader evolution in software engineering: the rising viability of building bespoke infrastructure for specific workloads using AI-assisted development tools. However, the deployment also underscores the delicate balance required when integrating artificial intelligence into critical systems, maintaining strict human oversight to safeguard production environments against potential automation failures.
Background Context and Architectural Bottlenecks
To understand the necessity of CobbleDB, one must examine the specific mechanics of Perplexity’s search architecture. Each user query triggers a complex retrieval pipeline. The system’s serving layer must rapidly pull pre-chunked textual passages and corresponding vector embeddings from a backend repository. During a standard Search API call, the system retrieves between 100 and 120 page keys in efficient batches of 10 to 20 items, with individual records averaging roughly 50 kilobytes in size.
Initially, Perplexity relied on Amazon DynamoDB to manage these workloads. While DynamoDB is a robust, globally recognized managed NoSQL database, it presented architectural and financial hurdles that grew more pronounced as Perplexity’s user base and corpus expanded. First, the managed service offered limited control over low-level read performance configurations. In a high-throughput search environment where responses must be aggregated instantly, a single lagging read replica can stall an entire batch request, degrading the user experience.
Second, the economics of cloud scalability began to shift unfavorably. Perplexity’s operations generate a continuous, high-volume stream of large reads and writes driven by concurrent search requests, web crawling initiatives, and ongoing document reprocessing. Under traditional managed pricing models, these sustained throughput levels translated to escalating cloud infrastructure costs that became increasingly difficult to justify. These limitations forced Perplexity to rethink its architecture, leading to the decision to decouple long-term document storage from the database dedicated to serving live, real-time searches.
The Chronology of Development
The creation of CobbleDB was characterized by rapid iteration, disciplined human architecture, and heavy reliance on artificial intelligence tools. The timeline of this infrastructure pivot unfolded over several distinct phases:
Phase 1: Diagnosis and Decoupling (Early Stages)
Perplexity engineers identified that DynamoDB could not cost-effectively sustain the required read performance thresholds for the platform’s scaling search traffic. The team made the strategic decision to architect a new storage stack divided into three discrete tiers to isolate long-term storage from live query processing.
Phase 2: AI-Assisted Implementation (Two-Month Sprint)
Tasked with building a purpose-built key-value store, a core team of two engineers leveraged hundreds of persistent coding agents to accelerate development. Operating within an intensive two-month window, the human-AI hybrid team wrote approximately 40,000 lines of Rust code. The coding agents maintained context across multiple sessions, proactively identifying runtime configuration edge cases, resolving restore assumption flaws, and generating rigorous unit and integration tests.
Phase 3: Deployment and Cutover
Following initial testing and optimization, CobbleDB was integrated into Perplexity’s production environment to handle live search queries. Comparative performance metrics were captured immediately before and after the cutover, establishing significant latency and throughput benchmarks.
Phase 4: Future Roadmap and Open-Sourcing
With CobbleDB successfully integrated into production, Perplexity has announced plans to eventually open-source the database, allowing the broader developer community to benefit from and contribute to the custom Rust storage engine.
A Three-Tiered Architecture for Modern Search Data
To achieve its performance goals, Perplexity engineered a sophisticated three-tier storage stack designed to optimize data ingestion, durability, and low-latency retrieval.
The foundation of the architecture is Pillar, which maintains durable document states on high-capacity hard disk drives (HDDs) utilizing YTsaurus. Pillar handles the heavy lifting of storing versioned metadata, document chunks, and vector embeddings over the long term.
The second tier acts as a distribution bridge. A component designated as Lorry packages updates into partition-specific batches and routes them securely through Amazon S3 directly to CobbleDB.
The final serving tier, CobbleDB itself, distributes processed page data across three independent replicas per data partition. Using hashed URLs as storage keys, the database leverages RocksDB to retain frequently accessed, hot data directly in memory, while warm and cold data reside on local NVMe storage. Read operations are optimized to remain within the same availability zone whenever possible. Crucially, if the system detects a slow replica during a batch request, the query router can dynamically reroute the read request to an alternative replica, preventing a single slow node from bottlenecking the entire batch. Furthermore, updates delivered via S3 are applied independently across nodes, allowing an individual replica to temporarily fall behind and catch up asynchronously without blocking cluster operations.
Supporting Data and Performance Benchmarks
When Perplexity evaluated its infrastructure performance following the cutover to CobbleDB, the metrics revealed substantial operational gains. At the time of measurement, Perplexity was handling approximately 200,000 requests per second.
Under these production conditions, the median batch-read latency for CobbleDB dropped to 5.6 milliseconds, representing a roughly fivefold improvement over the 31.4 milliseconds recorded under DynamoDB. The tail latency metrics showed an even more pronounced enhancement: the 99th percentile (p99) latency plummeted from 123 milliseconds on DynamoDB down to 24.2 milliseconds on CobbleDB. Subsequent stress testing and synthetic load benchmarks pushed CobbleDB even further, with the datastore sustaining up to 500,000 requests per second before encountering performance degradation.
From a financial perspective, Perplexity’s internal cost models indicate that CobbleDB will reduce database-related expenditures by at least 20 percent compared to optimized DynamoDB commitment options. However, independent industry analysts note that these financial projections primarily reflect direct cloud infrastructure savings and do not fully account for the hidden, long-term engineering overhead required to maintain, patch, and support a custom datastore in-house. Furthermore, comparative benchmarks were recorded across different timeframes—DynamoDB metrics prior to the cutover and CobbleDB metrics afterward—rather than through controlled, simultaneous side-by-side testing against identical traffic profiles.
Industry Perspectives and Governance
The deployment of CobbleDB touches upon a wider debate within the enterprise technology sector regarding the boundaries of artificial intelligence automation. Earlier this year at Percona Live, Carnegie Mellon University professor Andy Pavlo delivered a keynote warning that databases represent the most difficult and critical challenge for autonomous AI agents. Pavlo emphasized that mistakes made by automated agents involving production data can often be catastrophic, difficult to detect, and impossible to reverse.
Perplexity navigated this risk by enforcing a strict division of labor between artificial intelligence and human engineers. While hundreds of AI coding agents were permitted to write code, manage session contexts, and assist with testing, they were strictly barred from accessing or operating production database environments. The two human engineers retained absolute control over the system architecture, code review pipelines, and production deployments.
This hybrid approach mirrors trends observed across other technology leaders, such as Shopify and Ramp, which have successfully utilized custom coding agents built around third-party large language models to accelerate software delivery while maintaining rigorous human governance over core infrastructure.
Broader Impact and Implications
Perplexity’s successful transition from a managed cloud service to a custom-built, AI-assisted storage engine illustrates a shifting economic calculation in modern software engineering. Traditionally, building a custom database from scratch required large, specialized teams of systems engineers and substantial capital investments, making managed services the default choice for fast-growing startups.
CobbleDB demonstrates how AI-assisted development tools lower the barrier to entry for building complex, performance-critical infrastructure, empowering smaller engineering teams to tailor their technology stacks directly to their unique workload requirements. By reclaiming control over read performance and eliminating the marginal costs associated with high-frequency cloud queries, Perplexity has solved an immediate scaling bottleneck.
At the same time, the project serves as a case study in the trade-offs of infrastructure ownership. While custom datastores can yield superior performance and long-term cost efficiencies, they trade financial operational expenses for engineering maintenance debt. As Perplexity prepares to eventually open-source CobbleDB, the broader software development community will closely watch whether the performance gains and cost savings of AI-built databases outweigh the ongoing responsibilities of custom infrastructure ownership.
