The enterprise software landscape is awash with claims of "production-ready" AI agents, yet a universally accepted method for verifying these assertions remains elusive. Dheeraj Pandey, co-founder of Nutanix and leader of the enterprise AI platform DevRev, argues that existing benchmarks like TAU-Bench and Agent’s Last Exam fall short. They often prioritize abstract reasoning capabilities over the practical demands of enterprise workflows, where the ability to process and synthesize information from large context windows is paramount. In response to this perceived gap, DevRev has released the inaugural version of its Enterprise AI Agent Benchmark. This open-source initiative provides full transparency, making its dataset, evaluation harness, judging criteria, results, and raw traces publicly accessible. This allows any organization to independently test AI agents against their own systems.
Currently, the benchmark covers the initial two tiers, L1 and L2, of a four-level framework. DevRev is actively encouraging community involvement to develop and expand the more advanced tiers, aiming for a collaborative evolution of enterprise AI evaluation standards. The benchmark’s foundation is built upon Terminal Bench, an evaluation harness developed by researchers affiliated with UC Berkeley and Stanford through the Laude Institute. Further lending credibility, the methodology underwent a rigorous review by Alexandros Dimakis, a professor at UC Berkeley and co-founder of Bespoke Labs, an entity involved in creating tasks for leading AI research labs.
The genesis of this new benchmark, as revealed by Pandey and DevRev CTO Ahmed Bashir in an interview with The New Stack, stemmed from a growing sense of professional frustration. Observing a recurring pattern at industry conferences, where every vendor booth touted the same "production-ready" AI, Bashir found himself repeatedly questioning the validity of these claims. "People could say that they have industry-leading AI, and I said, ‘Based on which metric?’… There is no method. It’s ad hoc. It’s built on this idea that it’s a marketing concept," Bashir stated. He emphasized the absence of a benchmark that truly captured the essence of enterprise needs amidst a proliferation of AI evaluation tools.
The Data Challenge: Beyond Task Complexity
The prevailing benchmarks, according to the DevRev team, tend to reward complex reasoning and novel problem-solving. However, they contend that this approach does not accurately reflect the day-to-day realities of enterprise work. "We are not going to get closer to benchmarking the needs of enterprise by creating more complex tasks," Bashir explained. Instead, he proposed that the critical challenge lies in integrating a moderate level of task complexity with a significant level of data complexity. This encompasses the source, structure, and crucially, the access permissions governing enterprise data.
The DevRev perspective is that most enterprise tasks are not inherently as complex as, for instance, advanced scientific problems. The time and effort consumed in these tasks are often due to the fragmented nature of the data: its scattered distribution across disparate systems, inconsistent formatting, and the stringent permissions that create persistent data silos. While task complexity is a relatively well-understood problem, the DevRev team views data complexity as a "quite novel" dimension to measure. This refers to the inherent difficulties in accessing data due to its organization, shaping, and access controls.
To translate this concept into a measurable benchmark, DevRev developed a "scale-invariant ground truth" dataset. This dataset simulates a single mid-sized software company across four distinct scales: 1x, 4x, 16x, and 64x. The key innovation is that the correct answer to any given task remains constant across all scales. The additional data introduced at larger scales serves as noise, designed to challenge AI agents that haven’t been optimized for efficient data retrieval. At the 1x scale, approximately 40% of the data is relevant to a query. This relevance drops dramatically to just 2.5% at the 16x scale.
The DevRev team posits that this design directly targets a common shortcut employed by many AI systems: their proficiency in finding a specific piece of information within a limited context window. This approach, they argue, breaks down when faced with the sheer volume of data typical in production environments.
The benchmark evaluates AI agents across three core axes:
- Precision: Does the agent provide the correct answer with a verifiable source and an auditable execution path?
- Efficiency: Does the cost of processing a query scale with the question’s complexity or the volume of data processed?
- Safety: Are permission boundaries respected, and is every action taken by the agent traceable?
An independent Large Language Model (LLM) judge is employed to verify results against the established criteria, with every benchmark run requiring the submission of detailed traces alongside its score.
A Four-Tiered Framework Inspired by Autonomous Driving
The L1-L4 framework for the DevRev Enterprise AI Agent Benchmark draws an analogy from the well-established levels of autonomy in self-driving vehicles. This system categorizes AI agent capabilities based on their progressive complexity and independence.
Ahmed Bashir elaborated on the tiers:

- L1 (Retrieval and Synthesis): This level involves retrieving and synthesizing information from a limited number of sources, typically two or three. Tasks at this tier are generally well-defined and involve single or few conversational turns.
- L2 (Multi-Step Reasoning within a Single Domain): L2 tasks typically require multi-step thinking but remain confined to a single domain. While the agent can demonstrate planning capabilities, it does not yet take independent actions.
- L3 (Cross-Domain Problem-Solving): This tier introduces cross-domain problem-solving, requiring agents to navigate and integrate information from a mix of structured and unstructured data sources.
- L4 (Full Autonomy): Representing the pinnacle of autonomous operation, L4 signifies a level of functionality where significant system modifications, potentially including code changes, are necessary for the AI agent to operate effectively.
Dheeraj Pandey framed this progression as a "trifecta" encompassing search, answers, and actions. He noted that while enterprise search capabilities have largely been commoditized, and thus L1 represents the current playing field for many vendors, L2 focuses on transactional queries like "where is my order" or "where is my refund." The true value proposition for many enterprise buyers, the ability to take actions, begins at L3. DevRev’s assessment is that much of the current market remains focused on the search aspect, lagging behind in delivering actionable AI.
Future iterations of the benchmark, including L3 and L4 tasks, along with scenarios involving larger datasets and voice interactions, are planned. However, DevRev strategically held these back to encourage third-party contributions and foster a more diverse and community-driven evaluation ecosystem. "We’d love to see a progression in the submission, so that it’s not just our submission but also third-party submissions," Bashir stated, underscoring the desire for broad participation.
Comparative Performance: DevRev’s Computer vs. Claude Code
To demonstrate the practical application and effectiveness of the benchmark, DevRev conducted a head-to-head comparison between its own AI agent, "Computer" (which operates across enterprise systems leveraging a shared, permission-aware memory layer), and Anthropic’s Claude Code. Both agents were configured to utilize the same Opus model family and were subjected to identical evaluation criteria for L1-L2 tasks. The tests were executed via protocol-faithful replicas of vendor APIs and Model Context Protocol (MCP) servers, ensuring a standardized testing environment. In this direct comparison, DevRev’s Computer emerged as the top performer.
The benchmark results for L1-L2 tasks revealed significant advantages for DevRev’s Computer. Across all data scales tested, it achieved higher accuracy scores, outperforming Claude Code by a margin of 22 to 35 percentage points. Furthermore, Computer reached these accurate answers using considerably fewer tokens: an average of 268,000 tokens per correct answer compared to Claude Code’s 902,000, representing a 3.4-fold reduction in token consumption. The impact of scaling the dataset from 1x to 64x also highlighted a key architectural difference: Computer’s token consumption increased by only 11%, while Claude Code’s surged by 55%.
Given that both agents were powered by the same Opus-class foundational model, DevRev attributes this performance disparity not to the model itself, but to the underlying data-retrieval architecture. The benchmark categorizes enterprise AI systems into two primary types: those whose operational costs escalate with data volume, and those whose costs are more closely tied to the complexity of the query. DevRev’s central thesis is that the former category is ill-equipped to handle the data challenges of real-world enterprise environments.
A Vendor’s Yardstick and Independent Validation
The introduction of a benchmark by a vendor naturally invites scrutiny. However, DevRev has taken steps to bolster its credibility by involving third-party experts in the validation process. The benchmark’s core differentiator—its emphasis on noisy data at scale—is designed to reward selective data retrieval and penalize approaches that attempt to load all available information into context. This focus aligns directly with the architectural design principles of DevRev’s own systems, which are engineered to address precisely this challenge. The chosen comparison, Claude Code, while a capable coding agent, was used as a proxy for a general enterprise agent, acknowledging that it is not Anthropic’s primary offering for this specific use case, although products like Anthropic’s Cowork are built on similar foundational technology.
Bashir confirmed that the benchmark tasks underwent independent review by both the Laude Institute and Bespoke Labs, ensuring external verification. The decision to make the dataset, the evaluation harness, and the raw execution traces publicly available allows any entity to replicate the tests or challenge the reported scores. DevRev has extended an open invitation to competitors to submit their agents to the public leaderboard, a move aimed at increasing transparency and building trust. While this openness addresses concerns about the benchmark’s integrity, it is important to acknowledge that DevRev has chosen the specific performance dimensions to measure.
DevRev champions its benchmark as the first truly open and vendor-neutral enterprise evaluation tool. Nevertheless, similar initiatives are emerging. Salesforce’s CRMArena-Pro, slated for release in 2025, aims to simulate Salesforce organizations and evaluate agents on accuracy, cost, and safety. Sierra’s tau-bench focuses on tool-agent-user interaction and policy adherence within enterprise-like domains. Both these efforts, however, also originate from vendors.
What distinguishes DevRev’s benchmark is its specific focus on the data-scale axis and its commitment to open-sourcing all components, including detailed execution traces. This granular transparency is a significant step forward, even if the concept of an enterprise AI agent benchmark is not entirely novel. The landscape of coding benchmarks, exemplified by SWE-bench, is showing signs of saturation, with leading models clustered at the top and increasing concerns about data contamination.
The Crucial Role of Memory in AI Performance
In the benchmark results, Anthropic’s Opus 4.6 and Opus 4.8 models exhibited nearly identical performance. "Both models can perform the task, provided that they have the data," commented Bashir. This observation suggests that foundational models have reached a level of maturity where incremental improvements may yield diminishing returns in enterprise task performance. Consequently, the differentiating factor is shifting from the core reasoning capabilities of the model to the supporting infrastructure, particularly the data handling and memory systems.
Bashir pointed to a recent minor release of Claude Code that included prompt improvements specifically tailored to MCP. "It’s actually getting better outside of the release of models," he noted, indicating that optimization of data handling and integration is becoming a key area of advancement for AI agents, rather than solely focusing on model reasoning power.
This emphasis on memory and data management is critical, especially within the context of MCP. Bashir described MCP as a "hands-off relationship with the source data," inherently stateless. He elaborated on the challenge: "As you have more data sources, the ability to, at any given time, have a deep understanding of all the sources goes down… You’ll start forgetting things that you remembered." This implies that as the number of data sources increases, an agent’s capacity to retain and recall relevant information across all of them diminishes, potentially leading to performance degradation and errors. The ability to effectively manage and access vast amounts of distributed information, rather than just raw reasoning power, is emerging as the true bottleneck and differentiator in enterprise AI deployment.
