Breaking Through Visual Large Model Bottlenecks, New Intelligent Vision Chip Directly Translates Optical Signals into Tokens
2026-08-20 15:08
Favorite

How to efficiently and with low power consumption achieve synchronized acquisition and tokenization of visual information is a critical challenge that urgently needs to be addressed in the field of physical artificial intelligence (AI) hardware. Recently, a team led by Professor Miao Feng and Associate Professor Liang Shijun from Nanjing University, in collaboration with a team from the National University of Singapore, has taken the international lead in developing an ultra-low-power intelligent vision sensor chip, named "OptiToken," that can directly convert light into tokens. The tokens generated by this chip can be directly fed into an encoder to achieve image recognition. The related research findings were published online on the 19th in the international academic journal *Nature Sensing*.

The autonomous perception, reasoning, and decision-making capabilities of physical AI in complex environments heavily depend on the ability of edge devices to efficiently run visual large models. However, edge devices, constrained by power consumption and computational power, struggle to perform real-time visual cognition locally.

"In the traditional visual perception pipeline, optical signals must undergo multiple stages—such as image sensor capture, analog-to-digital conversion, buffer transfer, digital patching, and embedding—before generating tokens that models can process. The frequent transfer of massive redundant data leads to persistently high energy consumption at the edge," Miao Feng explained. "We aim to complete tokenization directly within the sensor, bypassing this pathway at the physical level."

In this study, the team took the international lead by proposing to move the token generation process forward into the sensor, directly accomplishing functions such as image capture, image patching, and image patch embedding within the sensor itself.

Liang Shijun stated that the "OptiToken" chip developed by the team consists of a photosensitive memory array and peripheral addressing circuits. The team first constructed a photosensitive memory array based on monolayer molybdenum disulfide floating-gate phototransistors. Each pixel in the array integrates three functions: light sensing, storage, and analog computation. After optical signals are written in, they are stored in situ in the form of floating-gate charges. The peripheral addressing circuits can selectively activate pixel regions to complete image patching. For the selected image patches, the chip further applies a specific voltage sequence, enabling the stored optical information in the patch and the embedding matrix to perform multiply-accumulate operations in the analog domain, with the output current serving as the token.

"We have achieved physical computation on the chip where 'light goes in, and tokens come directly out,' fundamentally eliminating data transfer—the largest source of energy consumption," Liang Shijun explained. This "physical tokenization" approach transforms the chip from a passive image recorder into an active visual semantic generator. The vision transformer architecture system based on "OptiToken" achieves an image recognition accuracy of 87.3%, approaching the baseline of traditional software-based recognition, while improving energy efficiency in the tokenization stage by more than 10 times compared to conventional solutions.

Miao Feng noted that this achievement fundamentally overturns the serial paradigm of "sense first, buffer next, compute last," opening a new dimension for breaking through the computational power and power consumption ceilings of edge AI in the post-Moore era. This technology is expected to provide a high-efficiency, near-zero-latency visual cognition foundation for fields such as embodied intelligence, low-altitude drones, and intelligent security, accelerating the large-scale deployment of physical AI.

This bulletin is compiled and reposted from information of global Internet and strategic partners, aiming to provide communication for readers. If there is any infringement or other issues, please inform us in time. We will make modifications or deletions accordingly. Unauthorized reproduction of this article is strictly prohibited. Email: news@wedoany.com