Publication date: 24th July 2026
High-performance computing consumes enormous amounts of energy as AI and machine learning become more widespread. Therefore, there is a strong demand for technologies that can improve both computational performance and energy efficiency. Compute-in-memory (CIM) is a promising approach for reducing energy consumption by performing computation directly within memory. In this presentation, we introduce our previous work on CIM and discuss its relationship to inference and generation in machine learning. In our previous report, we showed that in-memory computing using 3D flash memory can be achieved with extremely low power consumption. Since the high input voltage, Vpp, for the word line (WL) driver circuit is externally supplied, both Vread, the WL voltage for unselected cells, and Vcc, the supply voltage for circuits other than the WL driver, can be reduced. As a result, the power consumption during read operation of single-level cells (SLCs) is reduced by 56% at Vcc and by 98% at Vpp. This significant reduction in memory access energy makes it possible to increase the number of activation blocks and WLs used for in-memory computing. We also propose a novel approximate search method that combines sequential multi-block activation with a current control cell (CC cell). Key vector data are stored across multiple blocks, while query vector data are applied as WL voltages, and the on-current of each bit line (BL) is determined by the CC cell. This method requires no additional WL control circuit and uses only one cell per data unit. We demonstrate that the inner product (IP) between key and query vector data can be determined with sufficient accuracy not only for 8-dimensional vectors but also for 128-dimensional vectors. As a result, energy consumption is reduced by 99.4%, and memory access energy reaches as low as 0.17 pJ/bit. This technology is fully compatible with conventional 3D flash memory and is essential for realizing energy-efficient in-memory computing.
The authors wish to thank A. Kawasumi, Y. Komano and S. Sasaki for the contributions and advice on the approximate nearest neighbor search.
