The publication of this technical paper by researchers from Carnegie Mellon University (CMU) and the University of California, Los Angeles (UCLA), marks a significant milestone in the field of automated hardware design. By leveraging the reasoning capabilities of Large Language Models (LLMs) within a structured multi-agent framework, the research team has addressed one of the most persistent bottlenecks in High-Level Synthesis (HLS): the translation of arbitrary software code into efficient, hardware-compatible architectures. The system, known as AgRefactor, demonstrates a 6.51× geometric mean speedup over existing state-of-the-art pragma tuning tools, signaling a shift from simple optimization to deep, structural code refactoring driven by artificial intelligence.
The Evolution of High-Level Synthesis and the Compatibility Gap
For decades, the semiconductor industry has sought to bridge the gap between high-level software development and low-level hardware description languages (HDLs) like Verilog and VHDL. High-Level Synthesis was introduced as the solution, promising to allow developers to write in familiar languages such as C or C++ while automated compilers generated the corresponding hardware circuitry for Field-Programmable Gate Arrays (FPGAs) and Application-Specific Integrated Circuits (ASICs).
However, the "HLS gap" remained a formidable obstacle. Most C/C++ code written for General-Purpose Processors (CPUs) is fundamentally incompatible with hardware synthesis. Software often relies on dynamic memory allocation, complex pointer arithmetic, and non-deterministic control flows—elements that do not map directly to the fixed, parallel logic gates of a hardware chip. Traditionally, transforming this software into "HLS-compatible" code required human hardware experts to manually refactor the source code, a process that is time-consuming, error-prone, and expensive.
Previous attempts to automate this process focused primarily on "pragma tuning." Pragmas are compiler directives that tell the HLS tool how to handle specific loops or arrays (e.g., unrolling a loop or partitioning memory). While effective for optimization, pragma tuning cannot fix fundamentally incompatible code structures. This is where AgRefactor departs from traditional Electronic Design Automation (EDA) tools.
Architectural Overview: The Agentic Workflow
The core innovation of AgRefactor lies in its "agentic" architecture. Rather than treating an LLM as a simple code generator, the researchers at CMU and UCLA designed a multi-agent system where different AI entities play specialized roles in the refactoring pipeline. This approach mimics the iterative workflow of a human engineering team.
- The Architect Agent: This agent analyzes the initial software code to identify hardware-unfriendly constructs. It evaluates the dataflow and memory dependencies, determining which sections of the code require structural changes rather than mere optimization.
- The Refactor Agent: Based on the Architect’s analysis, this agent performs the actual code transformations. It replaces dynamic pointers with fixed-size arrays, converts recursive functions into iterative loops, and restructures data types to fit hardware bit-widths.
- The Verification and Feedback Agent: This agent interfaces directly with commercial HLS tools (such as AMD’s Vitis HLS). It attempts to compile the refactored code and captures error logs, resource utilization reports, and timing estimates.
- The Evolver Agent: This is the "self-evolving" component of the workflow. If the synthesis fails or performance is sub-optimal, the Evolver Agent analyzes the failure, updates the system’s internal knowledge base, and refines the prompts and strategies used by the other agents for the next iteration.
This closed-loop system allows AgRefactor to learn from its own mistakes. By iteratively refining the code based on actual hardware synthesis results, the system eventually arrives at a solution that is not only compatible with the hardware but also highly optimized for performance.
Quantitative Analysis: Performance and Speedup Results
The effectiveness of AgRefactor was validated through a series of rigorous benchmarks, comparing its output against both unoptimized code and code optimized by current state-of-the-art (SOTA) pragma tuning tools. The researchers reported a 6.51× geometric mean speedup across a variety of computational kernels.
The 6.51× figure is particularly notable because it represents an improvement in the quality of the hardware generated. In hardware design, speedup is often a result of better parallelism (loop unrolling and pipelining) and more efficient memory access patterns. Because AgRefactor can refactor the underlying code structure, it enables optimizations that pragma-only tools cannot reach. For instance, if a piece of code uses a linked list, a pragma tuner would struggle to optimize it for an FPGA. AgRefactor, however, can refactor that linked list into a hardware-friendly static array, allowing for massive parallel processing.
Data from the technical paper suggests that AgRefactor’s performance gains were most pronounced in complex algorithms involving irregular memory access and nested loops. In these scenarios, the "self-evolving" nature of the agents allowed the system to discover non-obvious hardware architectures that outperformed manual designs created by non-expert programmers.

Chronology of Development in AI-Driven EDA
The emergence of AgRefactor in June 2026 is the culmination of several years of accelerating research at the intersection of AI and hardware design. To understand the significance of this timeline, one must look at the progression of the field:
- 2020–2022: Initial experiments using LLMs (like early GPT models) for Verilog and C++ code generation. These models often produced syntactically correct code that failed to meet hardware timing constraints or synthesis requirements.
- 2023–2024: The rise of "AI for EDA." Major vendors like Cadence and Synopsys began integrating AI into their toolchains, primarily for place-and-route optimization and power analysis.
- 2025: Research shifts toward "Agentic" workflows. Researchers began to realize that a single LLM prompt was insufficient for complex engineering tasks. The concept of "Chain of Thought" and multi-agent collaboration gained traction.
- Early 2026: The CMU/UCLA collaboration focuses specifically on the "HLS Compatibility" problem, recognizing that the inability to refactor code was the primary barrier to FPGA adoption in data centers.
- June 2026: The publication of "AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance," providing a scalable framework for autonomous hardware-software co-design.
Industry Implications and Expert Reactions
The implications of AgRefactor extend beyond academic research, touching on the broader semiconductor industry and the future of data center acceleration.
Democratization of Hardware Design:
By automating the refactoring process, AgRefactor lowers the barrier to entry for FPGA development. Software engineers who lack deep knowledge of hardware architecture can now deploy their algorithms on specialized accelerators with performance levels previously reserved for hardware experts. This could lead to a surge in custom hardware acceleration for niche AI models, bioinformatics, and financial simulations.
Reduction in Time-to-Market:
Manual refactoring is often the longest phase of the HLS design cycle. Industry analysts suggest that an agentic workflow could reduce the design-to-prototype phase from weeks to days. For companies in the fast-moving AI sector, this speed is a critical competitive advantage.
Impact on EDA Tool Providers:
While AgRefactor is a research project, its methodology is likely to be integrated into commercial EDA suites. Major players in the HLS market, such as AMD (Vitis) and Siemens (Catapult), are under increasing pressure to provide "AI-first" interfaces. The 6.51× speedup reported by the CMU/UCLA team sets a new benchmark that commercial tools will be expected to meet.
Inferred Expert Perspectives:
While the paper represents the views of the authors—Yang Zou, Zijian Ding, Yizhou Sun, and the renowned Jason Cong—the broader community has long anticipated this shift. Jason Cong, a pioneer in HLS and a professor at UCLA, has frequently advocated for "customizable computing." The success of AgRefactor validates his long-standing vision that hardware should adapt to software, facilitated by intelligent automation layers.
Technical Challenges and Future Outlook
Despite its impressive performance, the researchers acknowledge that AgRefactor faces challenges common to LLM-based systems. One such challenge is "hallucination," where an agent might propose a code transformation that is logically sound but physically impossible to implement on a specific FPGA chip due to resource constraints.
Furthermore, the "self-evolving" process requires significant computational resources. Running multiple iterations of an HLS compiler, which itself is a heavy computational task, alongside multiple LLM agents, necessitates a robust backend infrastructure. However, the researchers argue that the one-time cost of autonomous refactoring is far lower than the recurring cost of human engineering hours and the opportunity cost of delayed product launches.
Looking forward, the research team suggests that the AgRefactor framework could be expanded to support "hardware-software co-evolution." In this scenario, the agents would not only refactor the software but also suggest modifications to the hardware architecture itself (e.g., adjusting the number of memory ports or the topology of the interconnect).
Conclusion
AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance represents a fundamental shift in how we approach the design of specialized computing systems. By moving beyond simple optimization and into the realm of autonomous structural refactoring, the CMU and UCLA team has provided a glimpse into a future where hardware design is as fluid and accessible as software development. The 6.51× speedup is a powerful testament to the potential of agentic AI, suggesting that the next generation of silicon will not just be designed by humans using tools, but by autonomous systems that understand the deep, intricate relationship between code and chips. As the industry moves toward 2027 and beyond, the principles laid out in this paper will likely become the standard for high-performance heterogeneous computing.
