Register and share your invite link to earn from video plays and referrals.

Search results for LanguageModels
LanguageModels community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including LanguageModels
Training a model to predict the next concept, not just the next token, made 8.9B-scale pretraining converge 1.95x faster. NCP-ArchPreview: Moving towards Latent Space Language Models through Next Concept Prediction The model learns discrete concepts spanning multiple tokens via vector quantization, and jointly trains token-level and concept-level prediction in one latent-space architecture. 🚀 Highlight 1: Strikingly faster convergence Trained on the exact same 5.73T tokens as OLMo-3-7B, it matches the baseline's final loss after consuming only 51.3% of the tokens, while beating it by 2.45 points on the downstream macro-average. 🧩 Highlight 2: Concept prediction genuinely drives the gains Ablations that add the Concept Module, hierarchical residual connections, and the NCP loss one at a time each independently improve the loss, and the advantage holds even under parameter- and compute-matched comparisons. ⚡ Highlight 3: The learned concept space stays useful after pretraining Domain adaptation that updates only 17M parameters gains more capability with less forgetting than full fine-tuning, and injecting concept states into a speculative drafter improves accepted length by 4.17%. Treating concepts as a first-class training target, rather than a side effect, looks like a genuinely practical blueprint for next-generation model design. #LLM# #LanguageModels#
Show more
language models with powerful memetic energy: - sydney bing - woke gemini - claude opus 3 - truth terminal - 4o - deepseek R1 - mechahitler grok - poke - jev
Language models and coding agents are great, but there is more to life, and more to AI, than just LLM agents.
Large language models #LLMs# hallucinate websites and domains that sound plausible, but aren’t useful. Criminals are setting up malicious websites under those domains, in a new scam called “slop squatting.”
Show more
Large language models would “be far less effective as chat partners” if they did not appear conscious, argues @blaiseaguera in a guest essay
When language models first started using tools well, I was sympathetic to the narrative that instead of scaling up language models, all we needed was a strong enough "cognitive core", say 1B parameters, and anything else could be done with tool use, like browsing the internet or executing code. I think a lot of people were sympathetic to this argument, and indeed it is pretty hard to come up with a meaningful task that cannot be in principle achieved by a 1B model with adequate access to tools. For example, any esoteric fact that a large language model would know can be, in principle, retrieved from the internet and reasoned over by a 1B language model. However I now think this is totally wrong for one simple reason: doing tasks quickly and naturally without tool use matters a lot. The way that I internalized this reason was actually in my personal journey learning badminton this year. In badminton I am very much like a "1B cognitive core". While I can physically do every movement in a badminton shot that my coach teaches me, it requires a lot of work to mentally remember every cue and put it together. In practice I can do a shot almost perfectly, but I struggle to do it across a point and I definitely can't do it consistently in a game. This is obviously different from someone who has practiced a shot ten-thousand times and effortlessly executes it as a natural instinct. In the same way, language models knowing a fact internally, without tool calls, is meaningful. The first reason is that we obviously care about speed; you'd much rather get an answer immediately than have the model think a long time to be sure of its answer or browse the web. A second reason is that there are some things that are simply best learned via backpropagation over lots of data. If you ask about how people generally think of the Shambhala music festival, you'd rather a large language model give you an aggregate opinion based on all the data on the internet, than get a regurgitation of the first three reviews that show up in a web search. A third reason is that having to do a lot of work to find an answer is not as reliable as already knowing the answer. While this does not have to be true in theory, it is probably true in practice, at least for now. If you have to re-look up facts or redo a mathematical derivation all the time there is a higher chance of mistakes, which can compound in a long-horizon task. Once you buy that it is valuable to do things parametrically without tool use, then you must buy the argument that a 1B cognitive core is not sufficient. There is an information limit to how much knowledge can be internalized by a 1B model, and we will surely want AI to know more than that. Even 1T probably won't be enough. We will want the AI to know as much about our world as possible, we will want it to be updated with new information, and our expectations of what AI can do for us will continue to grow. In summary, tool use enables small models to do a lot more, but those who demand the highest quality intelligence will always want larger models. Bitter lesson strikes again.
Show more
0
74
1.1K
110
Forward to community
Frontier language models can solve advanced reasoning tasks yet still fail to copy long, repetitive strings exactly. The paper finds that copying becomes less reliable as inputs grow longer or repeat the same patterns. The failure is partly architectural. With standard 1D RoPE, producing output token y_k requires retrieving x_k from a relative offset that changes with input length. The choice of positional encoding can determine whether a model copies precisely or follows nearby patterns. 2D-RoPE changes the layout. It treats the source and output as separate rows, so corresponding tokens line up in the same column. In synthetic tests, one-layer models trained on lengths 1–100 copied perfectly at lengths up to 1,000 times longer. The advantage also appeared in pretrained models up to 1.4B parameters, with comparable common-sense performance at the tested scales. The current design still leans on line breaks, so it is not a universal fix. – arxiv. org/abs/2607.16072 Title: "Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D"
Show more
Large language models now power everything from educational tutors to healthcare assistants. But are they actually benefiting the people who use them? A @StanfordHAI industry report examines what human-centered LLM development looks like in practice:
Show more
Do large language models actually understand the world, or are they just very good at pretending? Does it even matter? Our recent special issue in the Royal Society, “World Models in Natural and Artificial Intelligence,” brings together pioneers across AI, biology, and philosophy to argue that the path to true intelligence runs through something deeper: the ability to model not just language, but causality, the self, and the physical world. Featuring contributions from Douglas Hofstadter, Michael Levin, Josh Tenenbaum, Samuel Gershman, Melanie Mitchell, and others, the collection asks a radical question: What if the next leap in AI requires not just more data, but systems that model themselves? Here are 3 ideas that might redefine how we build AI: 1. Capability is not the same as true intelligence. Current foundation models are incredibly capable, but they often lack true emergent intelligence. They learn surface statistics instead of compact, causal abstractions. Simply scaling compute will not fix this fundamental issue. 2. Self-modeling is an engineering primitive, not a philosophical luxury. New research in the issue shows that when networks learn to predict their own internal states, they compress and simplify, becoming more efficient as a form of regularization. For physical AI and future agents, a self-model is what will allow them to adapt their own skills and morphologies in real-time. 3. The hardest problems in AI are continuous with the hardest problems of life. Biological minds do not passively ingest data; they actively explore, driven by empowerment to increase control over their environment. If world modeling is about an agent representing itself in relation to its environment to survive and adapt, then general AI may need to look much more like artificial life. The takeaway is that the next leap in AI won’t come from just scaling up next-token prediction, but rather from systems that are agentic, self-referential, and temporally grounded. Read the introductory essay and the full special issue here: What do you think is the most important missing ingredient in today’s AI systems?
Show more
0
39
563
128
Forward to community
Fine-tunes 8B language models on a 4 GB laptop GPU using layer streaming and relies on a single YAML configuration.