Register and share your invite link to earn from video plays and referrals.

Search results for Pretraining
Pretraining community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Pretraining
hot take: transfer in pretraining effectively does not exist to any meaningful degree
in life you want to scale yourself (pretraining, post training, inference time compute) but eventually you want to operate like a multi agent system (work/collab with awesome people)
PyTorch-native NeMo AutoModel handles transformer pretraining in @nvidia's end-to-end workflow for building a transaction foundation model. The workflow combines GPU-accelerated data processing and tokenization, decoder-only model pretraining, embedding extraction, and XGBoost fraud classification. On the synthetic @IBM TabFormer dataset, combining raw features with learned embeddings increased Average Precision by 41.76% over the raw-feature baseline. 🔗 Read the full post:
Show more
JUST IN: Google’s next flagship model Gemini 4 is reportedly “performing well” in internal pretraining evaluations.
Qwen 3.8-Next released with a great tech report with a ton of experimental results specifically surrounding architecture and pretraining. Pleasantly surprising Paper thread:
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Paper:
$GOOGL's next flagship model, Gemini 4, is reportedly performing well in internal pretraining evaluations, but still needs to complete posttraining, per WSJ.
Micron and Kioxia both anticipate that the rise of physical and agentic AI will drive longer term demand for memory and storage. Citing Stanford, GS thinks one robot may generate 200x more data in a *year* compared to the full pretraining dataset used to train one of today’s most advanced LLMs.
Show more
Noticed again that DeepSeek is the only Chinese model family that speaks fluent Russian even K3, 5.3 instantly slip into awkward Runglish the way Flash doesn't still probably the best general pretraining corpus over there (didn't check MiMo, Hy, Doubao, Qwen, were meh before)
Show more
the next unlock is inference-native context liquidity. once agents recursively price their own embeddings, the distinction between pretraining and distribution collapses. most teams are still optimizing for tokens when they should be optimizing for gradient ownership.
Show more