Researchers built an analog chip that runs LLM attention 100x faster than an H100 and uses 70,000x less power.
It's called GainCellAttention.
Every AI chip you use is built on an 80-year-old flaw called the von Neumann architecture.
the processor and the memory are physically separated.
To generate a single word, a GPU has to shuttle massive amounts of data back and forth across this divide. Over and over again.
Moving that data costs up to 10,000x more energy than actually doing the math.
But a team of researchers published a paper in Nature that completely destroys this bottleneck.
They built an analog in-memory computing architecture.
Instead of moving data from the memory to the processor, they put the processor inside the memory.
Using emerging hardware called "gain cells," the AI performs its most expensive calculation, the attention mechanism, directly inside the storage arrays.
The data never moves.
The compute happens exactly where the memory lives.
The result? A massive leap in energy efficiency and speed. It completely eliminates the latency of fetching data for every single token.
It is building AI chips that function exactly like the human brainwhere memory and computation are the exact same thing.
顯示更多