The evolution of generative AI from simple text completion to autonomous agency has necessitated a fundamental shift in how large language models (LLMs) interact with the physical and digital world. At the heart of this transition is the selection of an action primitive—the core architectural mechanism that allows a model to execute commands, retrieve data, or perform computations. While early agent frameworks relied almost exclusively on simple tool calling, the industry is increasingly moving toward sophisticated code execution environments. This transition is not merely a technical preference but a strategic response to the mounting costs of token consumption, latency in multi-step workflows, and the inherent difficulty models face when processing large datasets in their active context windows.
The Mechanism of Modern Agentic Action
To understand the divergence between these two primitives, one must look at how an LLM communicates its intent to external systems. Tool calling functions as a synchronous, iterative loop. When a model determines that an action is required, it emits a specific structured signal—typically a JSON payload containing the function name and its required arguments. The host application intercepts this signal, executes the underlying code, and returns the output to the model as if it were a new message in the conversation. This method is highly transparent; because every step is logged and re-injected into the context, it remains the gold standard for auditing and simple, single-shot requests.
In contrast, code execution represents a shift toward asynchronous, programmatic reasoning. Rather than issuing a single request, the model is provided with a sandbox environment where it can author and execute full scripts. Under the framework of Programmatic Tool Calling, developers define specific tools as "allowed callers," enabling the model to write code that iterates, performs conditional logic, and aggregates data without the host application needing to pause for a round-trip to the model for every intermediate calculation.
Chronology of the Shift
The industry’s movement toward these primitives can be traced through the release cycles of major AI research labs. In early 2024, the publication of the CodeAct paper by Wang et al. provided a critical academic foundation, demonstrating that agents utilizing executable code for multi-step tasks achieved a 20% higher success rate than those relying on traditional, prompt-based tool usage.
By late 2025, these findings were formalized into production-grade infrastructure. In November 2025, Anthropic introduced its Advanced Tool Use suite, which effectively bridged the gap between basic function calling and full code execution. This release allowed developers to mark specific tools as callable from within a model-generated script. This period marked a turning point: the industry moved away from treating code execution as a specialized, "hacky" workaround and toward viewing it as a standard feature of agentic architecture.
Quantitative Analysis of Performance and Cost
The implications of choosing one primitive over the other are best illustrated by the tangible performance metrics observed in enterprise deployments. For complex tasks involving "fan-out"—the process of querying multiple data sources and aggregating the results—the efficiency gains of code execution are substantial.
Anthropic’s internal benchmarking during the rollout of their programmatic tools revealed that on complex research-oriented tasks, token consumption dropped from an average of 43,588 tokens to 27,297. This represents a 37% reduction in processing overhead. More importantly, this was not a trade-off against quality. The same studies showed that accuracy on the GAIA benchmark—a rigorous standard for AI agent performance—increased from 46.5% to 51.2%.

Further evidence of this efficiency is seen in document-heavy workflows. In a case study involving the integration of Google Drive data into Salesforce CRM systems, migrating from standard tool calls to a code-execution-based approach reduced total token usage by 98.7%, from 150,000 tokens to 2,000. By keeping the bulk of the data within the sandbox environment and only returning the synthesized final output to the model, the architecture minimized the "context bloat" that previously drove up costs and degraded performance.
Comparative Framework: Selecting the Correct Primitive
For engineering teams, the decision to implement code execution versus tool calling should be dictated by the specific requirements of the task at hand. The following table provides a decision-making framework based on current industry best practices.
| Factor | Favors Tool Calling | Favors Code Execution |
|---|---|---|
| Call Frequency | Single or fixed, low-count calls | High-frequency or iterative calls |
| Data Handling | Requires human-readable reasoning | Requires aggregation or filtering |
| Data Sensitivity | Low; data is safe for context | High; keeps PII in the sandbox |
| System Latency | Low tolerance; direct execution | Tolerant; multi-step logic required |
| Infrastructure | Minimalist; no sandbox needed | Mature; requires secure isolation |
| Audit Needs | Individual action logging is critical | Final outcome is the priority |
Operational Implications and Challenges
While code execution offers superior performance for large-scale data tasks, it introduces new operational complexities. The most significant of these is the requirement for a secure, isolated sandbox. If an agent is allowed to write and execute code, the host environment must be hardened against potential injection attacks or malicious code execution. This necessitates the use of containerized environments or specialized runtime isolation, which adds a layer of infrastructure management that simple tool calling does not require.
Furthermore, debugging becomes inherently more complex. In a tool-calling architecture, the model’s intent is clearly articulated through JSON schemas, making it easy to trace exactly where a failure occurred. With code execution, the failure might be buried deep within a loop or a script written by the model, making it significantly harder for developers to reconstruct the decision-making process after the fact.
The Hybrid Future of Agentic Systems
Industry experts increasingly agree that the dichotomy between tool calling and code execution is a false one. Most robust production systems are designed as hybrid architectures. An agent may use traditional tool calling to check the current date or verify a user’s permissions—tasks that require simplicity and directness—and seamlessly switch to a code-execution environment to analyze thousands of rows of financial data or perform complex mathematical modeling.
The integration of tools like "Tool Search" and "Tool Use Examples" supports this hybrid model. These features allow models to navigate vast libraries of potential actions, selecting the right primitive based on the immediate context of the request. This capability is vital as agentic systems move into more complex domains, such as medical diagnostics, automated financial auditing, and large-scale supply chain management.
Broader Impact and Conclusion
The movement toward programmatic action primitives is a clear signal that the "AI Agent" category is maturing. We are moving past the era where agents were treated as conversational chatbots that occasionally "called" a function. Instead, they are becoming sophisticated software engineers in their own right, capable of designing their own workflows to solve problems efficiently.
The choice of action primitive is, ultimately, a choice about how much control the developer retains versus how much the model is empowered to orchestrate its own operations. By offloading logic to code, organizations can reduce costs, improve accuracy, and lower the latency of their AI-powered systems. However, this power comes with the responsibility to maintain secure, auditable, and robust infrastructure. As the technology continues to evolve, the distinction between "writing code" and "requesting a tool" will likely continue to blur, but the fundamental trade-offs in performance, cost, and complexity will remain the primary drivers of architectural design for the foreseeable future. The goal for developers is no longer to favor one method over the other, but to master the orchestration of both to build agents that are as performant as they are reliable.
