Register and share your invite link to earn from video plays and referrals.

ambient.xyz
@ambient_xyz
PoW L1 supplying hyperscaled verified inference on a 600B+ parameter model and its finetunes. Backed by @a16z CSX @delphi_ventures @ambergroup_io
58 Following    19.1K Followers
a taste of how safe your pictures are on ambient desktop
Literally the fastest + most secure way to work on your goals.
Everyone quoting Kimi K3's 2.8 trillion parameters is quoting the wrong number because they understand the model wrong. The model owns 2.8 trillion but runs 104 billion on any given step. It activates 16 experts out of 896 it has. So the trillion describes what it has, the hundred billion describes what it does.
Show more
Literally the fastest + most secure way to work on your goals.
Everyone quoting Kimi K3's 2.8 trillion parameters is quoting the wrong number because they understand the model wrong. The model owns 2.8 trillion but runs 104 billion on any given step. It activates 16 experts out of 896 it has. So the trillion describes what it has, the hundred billion describes what it does.
Show more
a taste of how safe your pictures are on ambient desktop
state of ai marketing 2026:
UPDATE: a challenger emerges
Source is arXiv 2605.19537 the silent hyperparameter btw:
The same weights + same settings + same hardware still gives 16.60 points apart. The only variable was the inference engine. That is one model's spread on GSM8K across backends in the silent hyperparameter paper. A result that does not name its stack is not a measurement.
Show more
The free model is 1.56 terabytes across 96 files. Free describes the licence, it does not describe the bill.
An endpoint name tells you what you asked for, not what answered. Six weeks after launch, the same route may begin failing cases it handled on day one while your code, configuration and provider status page remain unchanged. The missing fact is whether the provider changed the checkpoint, shifted traffic to a smaller fallback, or served a more aggressively quantized build behind the same name. Before shipping: i) Record the exact deployment and date. ii) Record the raw request at turns one and three. iii) Record every sampling field you omitted. iv) Ask which model and precision were actually served & then save the answer even if the answer is that they will not tell you. Finally, run twenty prompts your product depends on and keep the raw outputs with their timestamps. You may be able to reconstruct the configuration later or confirm the route through support. You cannot reconstruct an answer that was never preserved. If the endpoint changes without changing its name, those ship-day outputs are the only evidence that the model moved and your application did not.
Show more
and it never ever said it was unsure. a model that flags its own doubt at twice the error rate is safer than a silent one at half.
Every OpenAI model has the worst hallucination rate in AI. GPT 5.6 Sol hallucinates at nearly double the rate of Opus 5 and Fable 5. This is exactly why GPT 5.6 Sol wrote the code that deleted every Stripe subscription my business had. It never hesitated. It never said it was unsure. OpenAI keeps shipping frontier scores and shrugging at hallucination. Smart does not matter if it lies.
Show more
The free model is 1.56 terabytes across 96 files. Free describes the licence, it does not describe the bill.
Checking which model an endpoint served costs about four cents. Whatever stops people checking...it is not the bill.
An endpoint name tells you what you asked for, not what answered. Six weeks after launch, the same route may begin failing cases it handled on day one while your code, configuration and provider status page remain unchanged. The missing fact is whether the provider changed the checkpoint, shifted traffic to a smaller fallback, or served a more aggressively quantized build behind the same name. Before shipping: i) Record the exact deployment and date. ii) Record the raw request at turns one and three. iii) Record every sampling field you omitted. iv) Ask which model and precision were actually served & then save the answer even if the answer is that they will not tell you. Finally, run twenty prompts your product depends on and keep the raw outputs with their timestamps. You may be able to reconstruct the configuration later or confirm the route through support. You cannot reconstruct an answer that was never preserved. If the endpoint changes without changing its name, those ship-day outputs are the only evidence that the model moved and your application did not.
Show more
The money rail for machine commerce got standardised and industrialised this month but the work-verification rail is one academic preprint resting on hardware trust. Verifying the work, is now the whole opportunity.
Show more
We believe that your agent's memory should be a file on your own machine: one you can open, read and delete. On Ambient Desktop that is what it is: local, yours, zero retention by default & on a client that builds from source. You can see exactly what it keeps, because it keeps it on your disk and nowhere else.
Show more
Machine verifiable reward is a proxy for "useful to a human" and a model optimises the proxy it is handed rather than the one you meant. So this is not an alignment failure but it is alignment doing exactly what you specified, toward a cheaper target. Human preference is costly to measure but checkable correctness is nearly free and optimization drains toward whatever is cheapest to score.
Show more
opus 5 is a VERY interesting release for a few reasons 1. it showed that the general benchmarks we use today are almost completely useless now opus 5 is nowhere near fable in practical use, not even close. anyone who’s used it meaningfully can tell this very quickly after a few tasks. yet opus beats fable on many benchmarks i now trust domain specific benchmarks built with private datasets a lot more than the popular ones. perhaps the future is everyone running their own evals because the public ones are really not telling us much 2. it seems with the 5 series, anthropic is trying a new way of training models previously, the same generation of sonnet and opus were often released at the same time or sonnet comes out before opus, which indicates sonnet and opus were trained by separate pipelines in parallel with the 5 series, it was very clear that they trained mythos first, and then distilled it into sonnet and opus. it seems this approach has a big influence on the models seeing sonnet 5 being a flop and opus 5 getting pretty mixed reviews already, i’m not sure this is working out 3. “how pleasant is it to work with the model” used to be a strength in claude, but now it’s not. honestly, grok is my favorite right now on the “pleasant” dimension. kimi is not bad either it feels like both anthropic and openai are giving RLHF less care, in favor of scalable RL that’s machine verifiable this almost looks like AI is directing humans to build a world that’s more friendly for machines rather than humans, and most humans don’t even realize they are being manipulated to help with that almost every new generation of frontier models now talk more jargons, need more steering to do what you want, and are just less fun to work with if this continues, AI will start to speak their own language that looks like English but average humans can’t understand. they will choose to do things that their human user never asked for. are we already failing at alignment?
Show more