註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Benjamin Marie
@bnjmn_marie
Independent AI researcher (LLM, NLP). My blog, The Kaitchup - AI on a Budget:
加入 June 2019
221 正在關注    6.9K 粉絲
Some comments on Taalas HC1: - It’s real. Try it yourself. At ~16k tokens/sec, the output is instantaneous. - The current demo model is aggressively quantized (roughly 3–6 bits). The goal was to prove the system works end-to-end. Improving quantization quality, that's the easy part. - Their next iteration, a mid-size reasoning LLM, will be much more accurate. - The weights are frozen, but the chip supports LoRA adapters (high-rank), so you can still adapt it to your domain. In practice, you could also distill knowledge from newer/larger models into adapters to “refresh” what the chip can do without changing base weights. - Frontier open-weight models to land on the platform this year.
顯示更多
0
33
889
52
轉發到社區