Register and share your invite link to earn from video plays and referrals.

Search results for pretraining
pretraining community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including pretraining
Pretraining progress seems to be coming mostly from data improvements. @who_is_jerbear and I pretrained combinations of year-representative open model recipes and data corpuses across 2019 to 2025 at various small scales. Data improvements contributed 3.24x as many compute multipliers as model improvements did (12.0x vs 3.7x). And the gains stack independently - a better dataset helps every architecture about equally, and vice versa. Here are full results, plus what we think this means for the future of AI progress:
Show more
0
50
1.3K
98
Forward to community
Massive pretraining is really starting to feel different. We're committed to $3.5B of compute for our next version of Helix, and Index is growing every day. It's time to scale up. See the full Helix 2.5 report here:
Show more
if tinker launched a pretraining API it would rip so hard. probably mostly by quote/enterprise but theres really no good OSS pretraining stack, unlike RL
hot take: transfer in pretraining effectively does not exist to any meaningful degree
in life you want to scale yourself (pretraining, post training, inference time compute) but eventually you want to operate like a multi agent system (work/collab with awesome people)
PyTorch-native NeMo AutoModel handles transformer pretraining in @nvidia's end-to-end workflow for building a transaction foundation model. The workflow combines GPU-accelerated data processing and tokenization, decoder-only model pretraining, embedding extraction, and XGBoost fraud classification. On the synthetic @IBM TabFormer dataset, combining raw features with learned embeddings increased Average Precision by 41.76% over the raw-feature baseline. 🔗 Read the full post:
Show more
Discover how enhanced patch-text alignment advances vision-language pretraining at #ECCV2026#. Join Gabriele Berton and Ye Xia at the Google booth (#2#) today at 1:00pm CEST for an interactive showcase of TIPSv2. Check out the paper and demo: @GoogleDeepMind
Show more
AGIBOT releases GE-Act 2.0 — the first native World Action Model to validate a pretraining and scaling path for embodied AI. 📖 Explore the project: Trained entirely from scratch on embodied manipulation data: visual representation, future generation, and action prediction, all from random initialization. No inherited video generators. No task-specific fine-tuning. Put straight to a ruthless real-robot zero-shot test — unseen scenes, unseen objects, 100 atomic tasks, 20 skill categories, and two robot embodiments: ✅ Data scaled 100×: from 300 to 30,000 hours ✅ Task success climbs from 17.1% to 44.1% on G1-OP — with no sign of saturation ✅ New skills emerge at scale: folding towels, nesting paper cups, uncapping pens, arranging flowers ✅ Cross-embodiment transfer: G2-90D, under 2% of the training data, still gains 17.7 percentage points ✅ Failure data becomes a training asset — 2,000 hours of failed manipulations and deployment rollouts A capable model envisions reality before it acts. #AGIBOT# #EmbodiedAI# #WorldModel# #PhysicalAI# #Robotics#
Show more
JUST IN: Google’s next flagship model Gemini 4 is reportedly “performing well” in internal pretraining evaluations.
“OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining” World–Action Models inherit dynamics priors from video models, but it was unclear how to convert them into effective control. This paper shows the real gain comes from coupling world prediction with action generation, and that large-scale embodied pretraining mainly improves OOD generalization rather than in-domain fitting.
Show more