Today, users have to trust that providers are running the precision they claim.
What if inference could be verified?
We tested Qwen3.5-35B on 200 HumanEval prompts:
• FP8 was close to full precision
• INT4 drifted more
The path to decentralized, verifiable AI starts here.