가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Sudo su
@sudoingX
GPU/local LLM. more RAM and OSS... everywhere
가입 August 2022
1.1K 팔로잉 중    36.3K 팬
bonsai 2 27b just built this from one paragraph of prompt in one shot, all of it out of a 5.9gb file on an rtx 3060 12gb. i did not expect frontend taste at this size. small models usually get the logic right and the layout wrong, this one got the layout right, and it thought for a long time to do it, 41k tokens over 46 minutes, the context ran out to 77k and it never lost the thread. this is a ternary compression of qwen 3.8 27b, 26 tok/s fresh on a five year old gaming gpu, 13 tok/s at 77k deep, the whole 262k window resident. for what it is, on this card, this is insane, and i cannot wait to run it through real agentic coding on hermes agent, the tool loop, the builds that break and have to recover. @PrismML keep going, this is the one that runs on the card people actually own.
더 보기