SK hynix Memristor AI Chip Shows 21.3 TOPS/W But Leaves Throughput Gap
SK hynix, TetraMem and USC researchers developed a 65 nm memristor-based in-memory computing chip for edge AI. The paper listed 21.3 TOPS/W at 100 MHz, while the public record still lacks full-chip saturated throughput.

SK hynix, TetraMem and researchers from the University of Southern California have built a 65 nm memristor-based in-memory computing system-on-chip that achieved 80.36% accuracy on a lightweight AI benchmark and 21.3 TOPS/W at 100 MHz.
The experimental result validates the architecture’s low-power approach, but it does not establish how much throughput the complete chip can sustain.
Architecture and method
The system targets edge devices running lightweight neural networks, where moving data between memory and processing logic can consume substantial power.
Its embedded RISC-V processor schedules workloads across 10 neural processing units (NPUs).
Nine NPUs use conventional 256 × 256 memristor crossbars for analog vector-matrix multiplication.
Each includes 256 8-bit digital-to-analog converters, 256 8-bit analog-to-digital converters and peripheral circuitry for reading, writing, programming and controlling the array.
A tenth NPU is dedicated to depthwise convolution, a core operation in models such as MobileNet that maps poorly onto conventional crossbars because it filters each channel independently and offers limited data reuse.
TetraMem’s depthwise design replaces straight selection lines with a zig-zag topology.
Its NPU contains eight specialized 252 × 28 crossbar blocks, whose diagonal lines activate 252 memory cells across 28 columns.
That arrangement allows 28 independent 3 × 3 convolutions to run in parallel while using the full array for weight storage.
SK hynix fabricated the memristor devices and integrated the resistive-switching cells above 65 nm CMOS circuitry using its back-end process.
MobileNet demonstration
The researchers tested a customized MobileNetV1Small network on the Visual Wake Words benchmark.
The network contained approximately 36,000 parameters.
Depthwise layers ran on the dedicated NPU, while pointwise layers used five of the standard NPUs.
Because the hardware performs unsigned analog vector-matrix multiplication, the inputs and weights were quantized to unsigned 8-bit values.
Individual memristors provided slightly more than 2 bits of effective precision, so a two-subarray compensation technique increased effective weight precision to roughly 4 bits.
The chip reached 80.36% end-to-end inference accuracy, matching the corresponding 4-bit software model.
Its reported peak throughput was 0.254 TOPS per NPU.
Energy efficiency reached 21.3 TOPS/W at 100 MHz and 11.9 TOPS/W at 400 MHz.
The joint paper says those figures compare favorably with published SRAM-based compute-in-memory accelerators and exceed the Nvidia A100’s INT8 energy efficiency by an order of magnitude.
Efficiency and performance limit
The demonstration used only six of the 10 NPUs: one depthwise unit and five standard units.
Four standard NPUs remained idle, so it did not show sustained throughput for a real network with the full chip operating simultaneously.
The often-cited 2.54 TOPS figure is therefore a theoretical extension of the per-NPU result across all 10 units.
The paper does not disclose whether the NPUs can be saturated together, leaving full-chip throughput and sustained real-network performance unverified.
The work demonstrates a functioning research prototype with strong efficiency, but not yet a validated commercial edge-AI processor.




















