Researchers at the National University of Singapore (NUS) have unveiled a breakthrough in semiconductor architecture designed to tackle the escalating computational demands of Large Language Models (LLMs). The project, detailed in a technical paper published in August 2026, introduces CHIPSMORE, a specialized hardware framework that merges compute-in-memory (CIM) and compute-in-interconnect technologies. By utilizing a heterogeneous approach to data processing, the NUS team aims to solve the "memory wall"—the bottleneck where data movement between memory and processors consumes more energy and time than the actual computation—that currently hampers the scalability of artificial intelligence.
The Technical Foundation of CHIPSMORE
At the core of the CHIPSMORE architecture is the integration of two distinct types of memory-based computing: resistive RAM analog compute-in-memory (RRAM-ACIM) and static RAM digital compute-in-memory (SRAM-DCIM). This duality allows the chiplet to be highly versatile, capable of handling both high-precision base-mode inference and the specialized requirements of low-rank adaptation (LoRA).
LoRA has become an industry standard for fine-tuning massive models, allowing developers to adapt pre-trained LLMs to specific tasks without retraining the entire parameter set. However, running these adapted layers alongside base models often leads to significant latency. The CHIPSMORE design utilizes a programmable Inter-PE computational network (IPCN) to bridge the gap between RRAM and SRAM modules. This interconnection layer is not merely a transport medium; it is a computational fabric that allows for real-time adjustments in how data is processed, enabling the hardware to switch between modes dynamically based on the specific workload requests.
Chronology of Memory-Centric Computing
The shift toward compute-in-memory (CIM) did not emerge in a vacuum. For decades, the von Neumann architecture—which strictly separates the central processing unit (CPU) from memory—served as the bedrock of computing. However, as AI models grew from millions of parameters to trillions, the energy cost of shuttling data across the physical distance between memory and processors became unsustainable.
- 2020–2022: Early industry experiments focused on digital CIM (SRAM) to improve energy efficiency for small-scale edge AI.
- 2023–2024: The industry saw a surge in RRAM research, which offered higher density and non-volatile storage, essential for keeping massive model weights locally on the chip.
- 2025: The introduction of multi-tenant LLM inference servers created a demand for hardware that could handle multiple requests of varying complexity simultaneously.
- August 2026: The publication of the CHIPSMORE paper represents a maturation of these concepts, shifting the focus from individual CIM cells to a complex, heterogeneous chiplet system.
By moving the computation into the interconnects—the very pathways that typically serve as passive lanes for data—the NUS researchers have effectively reduced the physical footprint of the inference engine, allowing for higher throughput in power-constrained environments such as data centers and edge-computing nodes.
Data-Driven Performance Metrics
The architectural efficiency of CHIPSMORE is supported by the need for multi-mode, multi-request handling. In modern data centers, a single LLM server is often tasked with processing concurrent requests from different users, each requiring different model configurations or LoRA adapters. Traditional architectures often suffer from "context switching" latency, where the GPU or TPU must flush and reload memory buffers to accommodate these disparate tasks.
The CHIPSMORE framework mitigates this by allowing the RRAM-ACIM to handle the bulk of the static, heavy-weight parameters, while the SRAM-DCIM provides the flexibility required for rapid, frequent updates associated with LoRA. The programmable IPCN acts as a traffic controller, ensuring that the appropriate computational resources are allocated to the correct memory bank without creating data stalls. Preliminary simulations referenced in the study suggest that this architecture reduces energy consumption by approximately 40% compared to traditional GPU-based inference setups, primarily due to the elimination of redundant data movement across the traditional memory bus.
Industry Implications and Market Context
The implications for the semiconductor industry are significant. As LLM developers continue to favor parameter-efficient fine-tuning, hardware that can accommodate dynamic model changes will become a primary competitive differentiator. Currently, the market is dominated by general-purpose GPUs, which, while powerful, are not optimized for the specific, recurring matrix-vector multiplications inherent in LLM inference.

Analysts suggest that the modular, chiplet-based approach of CHIPSMORE is particularly well-suited for the post-Moore’s Law era. As monolithic silicon die yields become more expensive and difficult to manufacture, moving toward a "chiplet" strategy—where smaller, specialized dies are stitched together on a single package—has become the standard for companies like Intel, AMD, and NVIDIA. The NUS research provides a blueprint for how these chiplets could be utilized specifically for AI acceleration, potentially opening the door for specialized "AI-first" silicon.
Expert Perspectives on Architectural Shifts
While the research is in the preprint stage, the academic community has noted the importance of the programmable interconnect. "The bottleneck in modern AI is rarely the raw FLOPs (floating-point operations) available," noted a research lead at a major semiconductor firm who reviewed the study. "It is the ability to move and format that data efficiently. By integrating the computational logic directly into the interconnects, the CHIPSMORE design effectively turns the entire chip into an active participant in the inference process, rather than a collection of silos."
However, hurdles remain. The integration of RRAM into commercial high-volume manufacturing (HVM) has historically been plagued by reliability issues and thermal management constraints. For CHIPSMORE to move from an academic breakthrough to a production-ready component, the researchers must demonstrate that the RRAM-ACIM can withstand the heat generated by sustained, high-speed LLM inference operations over thousands of hours of continuous use.
Broader Impact: The Path Toward Sustainable AI
The broader environmental and economic impact of such a design cannot be overstated. With data centers currently accounting for an increasing percentage of global electricity consumption, the drive toward "Green AI" is gaining momentum. If the CHIPSMORE architecture can achieve even half of its projected efficiency gains in a real-world server rack, it could significantly lower the operational costs of deploying LLMs.
Moreover, the versatility offered by the dual-mode architecture (Base + LoRA) could democratize access to AI. Smaller entities and startups, which rely on fine-tuning pre-trained models rather than building them from scratch, would benefit from hardware that reduces the cost-per-inference of LoRA. By lowering the barrier to entry, this technology could accelerate the development of specialized LLMs for medicine, law, and engineering, where high-precision, fine-tuned models are essential.
Future Research and Development Trajectories
Looking ahead, the NUS team indicates that the next phase of development will focus on scaling the IPCN to support larger clusters of chiplets. As the system scales, the complexity of the interconnect routing becomes an exponential challenge. The research team is currently investigating AI-driven placement and routing algorithms that can configure the IPCN on-the-fly, potentially allowing the hardware to "learn" the most efficient paths for specific classes of LLM workloads.
Furthermore, the integration of high-bandwidth memory (HBM) with the CHIPSMORE chiplet is a logical next step to ensure that the system does not become starved of data from external sources. The convergence of logic, memory, and interconnects into a single, cohesive package is the defining challenge of 2026’s semiconductor landscape, and the CHIPSMORE project serves as a critical milestone in that evolution.
Conclusion
The CHIPSMORE project represents a sophisticated response to the limitations of traditional AI hardware. By rethinking the physical arrangement of compute and memory and introducing a programmable interconnect layer, the researchers have proposed a design that addresses the core bottlenecks of contemporary LLM inference. While the transition from research paper to commercial silicon involves significant engineering risks, the design principles outlined by the NUS team align with the industry’s shift toward heterogeneous, chiplet-based AI architectures. As the demand for more efficient, agile, and high-performance AI hardware continues to grow, innovations like CHIPSMORE will be instrumental in shaping the next generation of computing infrastructure, potentially marking a shift away from general-purpose GPUs toward specialized, memory-centric systems that offer a more sustainable future for artificial intelligence.
