A National University of Singapore preprint pairs analog and digital compute in memory chiplets — small specialized chips that perform math inside their own memory arrays — to attack the data movement bottleneck behind multi user inference cost.
The reason an AI chatbot feels slow under heavy load isn't that the chips can't think. It's that they can't move data fast enough between processor and memory. Under multi-user load, memory bandwidth becomes the binding constraint, not raw compute, and that is why inference cost is so stubborn.
Inference is the act of running an already-trained model to produce an answer, and it is the work that dominates real-world AI serving bills. A large language model (the class of AI behind chatbots and assistants) holds billions of parameters and a per-user cache of recent conversation, the so-called KV cache, that grows with every request. Move that data between the processor and off-chip DRAM and you pay in time and energy on every token. Stack thousands of users on a single server and the memory subsystem becomes the bottleneck long before the math units do.
This is the problem an August 2026 preprint from the National University of Singapore, called CHIPSMORE, tries to address. Rather than build one monolithic die, the team stitches together small specialized chips called chiplets. Two of those chiplets carry memory arrays that do computation locally: one is analog, built on resistive-RAM (RRAM-ACIM, a non-volatile memory that can perform matrix multiplies in place), and one is digital, built on SRAM (SRAM-DCIM, a fast on-chip memory that can also compute). A third piece, a programmable interconnect the team calls IPCN, lets the chiplets pass data and partial results between each other without round-tripping through a host processor. The name stands for compute-in-interconnect-and-memory; the design is as much about routing as it is about the memory arrays themselves.
The pitch is that for multi-request serving, where many users hit a model at once, you want different memory technologies for different jobs. Analog RRAM is dense and efficient for the heavy matrix multiplies inside the model. Digital SRAM is faster and more predictable for the smaller, more varied operations the model performs around those multiplies, including serving users who have a personal adapter, a low-rank adaptation (LoRA) layered on top of the base model. Putting both kinds of memory on the same package, and routing data between them through the interconnect, lets the design match the work to the right memory instead of forcing everything through one type. The paper claims support for both base-model inference and LoRA inference on the same hardware, which is the concrete capability CHIPSMORE puts on the table.
Semiconductor Engineering's coverage tags the work as a play on the memory hierarchy and KV cache rather than raw compute throughput, which matches the framing. This is a data-movement design, not a throughput design, and the trade publication's adjacent items that day point at the same broader problem from different angles: an RPI/IBM paper on HBM ECC controller overhead and a NYCU/TSMC study on molybdenum-based photomask materials, both running on the same chip-architecture feed, both touching the cost of moving and staging data on advanced AI hardware even if they solve different sub-problems.
The honest caveat is also part of the story. CHIPSMORE is a preprint, not a peer-reviewed result, and the headline performance numbers live in the arXiv PDF rather than the abstract. NUS has not announced a fab partner, a tape-out, or a deployment timeline, so the supported read is narrower than "AI chip breakthrough." It is one academic data point in a wider industry direction, alongside HBM stacking, on-package memory, and compute-in-memory work at other labs, all aimed at the same bottleneck.
The memory wall is not going away. Model sizes keep growing, KV caches swell with conversation length, and multi-user serving is the default rather than the exception. The architectural bet the industry is making is that pushing the math into the memory array, instead of shuttling data back and forth, is the lever that will bend inference cost down at scale. CHIPSMORE is one way of testing that bet with heterogeneous chiplets, and the same-day cluster of trade coverage suggests it is not the only group doing so this month.