Register and share your invite link to earn from video plays and referrals.

Search results for SmallModels
SmallModels community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including SmallModels
"Bigger is better" may no longer be a given. This is an attempt to lift small models to frontier-level quality through training automation alone 🔬 Title: Tiny AutoScientist: Supersized Intelligence for Small Models URL: 🔬 Overview Tiny AutoScientist is an automated research system that automates the entire training and alignment process for the small models (roughly 0.8B-8B) commonly used in production, aiming to make them perform at frontier-level quality. ❓ Challenges Solved In production, latency, cost, and device constraints often push you toward small models. ・But small-model training is hyperparameter-sensitive, prone to overfitting, and tricky to handle ・So you're often forced into a painful choice between a small model that fits constraints or a large one with enough capability 💡 Methodology & How It Works ・It automatically co-optimizes your data and model-training recipes ・It self-improves both until quality converges on your objective ・It automates the full R&D loop once reserved for frontier labs, absorbing the hyperparameter sensitivity and overfitting that plague small-model training 📊 Experimental Results ・35% relative improvement over human-configured training ・Consistent gains across dataset sizes from 5K to 100K samples ・Works across multiple model architectures ・Delivers frontier-level performance in days instead of months 🌍 Use Cases It unlocks previously impractical use cases: edge deployment, on-device inference, latency-sensitive apps, and regulated industries with strict data boundaries. Since small-model tuning tends to be artisanal, automating it to beat human-configured runs carries real practical weight. #SmallModels# #AutoML#
Show more
Small models. Big cost savings. @relace_ai trains specialized models purpose-built for code generation, powered by the dedicated @ycombinator GPU clusters on Together AI.
Anthropic has no small models that are worth using right now. OpenAI has no large models that are worth using right now. Google has no models that are worth using right now.
0
346
7.4K
232
Forward to community
The model and harness are designed together because small models fail in harnesses built for frontier models. Portable is a minimal system prompt, skills that load on demand, connectors as compact CLI tools instead of MCP servers, self-verification, and an always-on sandbox.
Show more
Super happy to release SmolDataEnvs: 5,000 verifiable RL environment tasks for hill-climbing small models in code and data science by @adithya_s_k 100% open source: environments, evals, training!
Show more
In 6 months, 95% of tasks will be done by small 250B open source models The large models will perform very complex tasks including train the smaller models The majority of token usage will be small models
Show more
🚀 MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale How can we make small models stronger for on-device agent deployment? MERA uses stronger models to guide an iterative loop of RL/GRPO, skill learning, and router optimization. Student failures become verified demonstrations, reusable SkillBook procedures, and LoRA updates, helping the small model take on more work over time. 🔥 Results: Qwen2.5-Coder-1.5B: 28.7% → 49.7% coding pass Qwen3.5-2B on TAU-2: 14/35 → 18/35 Fine-tuned 2B matches an unadapted 4B model Don’t just route around small models. Evolve them. 📄 💻
Show more
🚀 MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale What if small models could improve themselves, not just be routed around? MERA evolves both small-model skills and weights, guided by stronger models. Student failures become verified teacher demonstrations, reusable SkillBook procedures, and LoRA updates. 🔥 Results: Qwen2.5-Coder-1.5B: 28.7% → 49.7% direct pass on coding tasks Qwen3.5-2B on TAU-2: 14/35 (40.0%) → 18/35 (51.4%) The adapted 2B model matches an unadapted 4B model Small models don’t just get routed around. They get better. 📄 Paper: 💻 Code:
Show more
Good Chinese openweight models will be optimized for Chinese hardware first. For DeepSeek V4s versions, Huawei Ascend are some of the only two stacks with optimized inference ready, alongside CUDA. The exceptions to this might be the small models that fit gaming GPUs.
Show more
Raleigh, Oct 17-18. Open models, 33B or smaller. Prizes include an RTX 5090. Small Models Hack with Red Hat and @NVIDIAAI. Fine-tune, build multi-agent systems, or go deep on tool-calling. Apply now:
Show more