๐คฏ This is like having Jev at home, a AI reflex system running locally.
Someone is turning Google's open DiffusionGemma 26B-A4B into a local version of the new โSystem Oneโ AI idea inside vLLM.
Remember Jev? It is a decision LLM. Instead of asking an LLM to generate paragraphs, System One models are designed to make FAST structured decisions:
๐ฏ yes / no
๐งญ route A / B / C
๐ ๏ธ which tool to call
๐จ severity 1โ4
๐ classify / score / choose
And DiffusionGemma has an important role because a normal autoregressive LLM takes Prompt โ token โ answer
DiffusionGemma is Prompt โ predefined answer slots โ fill multiple decisions in parallel
Because it operates over a token canvas with bidirectional attention instead of being forced to generate everything strictly left-to-right.
The vLLM patch lets it return
โ
bounded choices
โ
probabilities
โ
confidence / uncertainty
โ
multiple decisions simultaneously
And it runs locally.
On ONE DGX Spark the developer reports
โก 1 request: ~120 ms
๐ concurrency 32: ~54 requests/sec
๐ง 3 decisions/request
๐ฅ ~162 decisions/sec
Then the developer found each question can require only around 3 canvas tokens:
๐ข index
โ placeholder
โ๏ธ separator
Meaning as many as ~85 questions could fit into 1 diffusion canvas.
And subsequent optimization reportedly cut decision time another:
โก ~20โ40% with no change in decision quality.
The model stats ...
๐ง 25.2B total parameters
โก 3.8B active
๐พ NVIDIA NVFP4: ~18.9GB
๐ฎ Can fit on a 24GB NVIDIA GPU
๐ Open weights
โ ๏ธ vLLM PR #
57250# is still OPEN ๐ not merged.
And this is Jev-like functionality. It is NOT evidence that DiffusionGemma matches Jev's intelligence or calibration.
๐ /vllm-project/vllm/pull/57250