Register and share your invite link to earn from video plays and referrals.

Dmytro Dzhulgakov
@dzhulgakov
Co-founder and CTO @FireworksAI_HQ. PyTorch core maintainer. Previously FB Ads. Ex-Pro Competitive Programmer
795 Following    7K Followers
Running fast is not enough, you need fast AND correct An excellent addition from @ArtificialAnlys to make sure that the flashy speed numbers are backed by 100% matching accuracy
Show more
Fear not, @FireworksAI_HQ is perfectly legal and serves the best models on 🇺🇸 servers for your AI-ndependence
do you know what you pay for in agentic workloads? cached tokens! session with 50+ tool calls -> prompt is billed 50 times all providers give 1/5 cached discount for GLM-5.2 we at @FireworksAI_HQ dropped it to 1/10, matching GPT/Claude that's -40% typical savings, have fun!
Show more
you may have heard that glm-5.2 at 392 token/s is cool, how about 446 except… it’s all noise. Artificial Analysis picks median among 8 points/day so first point of the day can be way off looking at 3 day better is better, but best is to test your actual workload, any single benchmark is not going to be representative
Show more
DSpark from @deepseek_ai ingeniously integrates many speculative decoding ideas to achieve 1.5x to 5x higher throughput in a real production system Let's understand it with 10 ideas, starting from the very basics 🧵
Show more
0
13
969
119
Forward to community
you may have heard that glm-5.2 at 280 token/s is cool, how about 318 and we still have room to go
bit-equivalent on-policy rl for glm-5.2 has been achieved internally developing…