Register and share your invite link to earn from video plays and referrals.

JJ
@JosephJacks_
540 Following    45.9K Followers
The @OSSCapital portfolio is always cooking 🍳 … Congrats to @liquidai on the @CellCellPress launch! Congrats to @compai on the Series A launch! Congrats to @cap on the v0.6 launch! Congrats to @planepowers on the Agents launch! Congrats to @dubdotco on the marketplace launch! Congrats to @calcom on the v6.9 launch! Congrats to @W4Games on the Series B launch! Congrats to @drinkleisure on the launch with Target! Congrats to @RadicalNumerics on the Omnii bio AI model launch! Always be launching ♾️
Show more
PSA: @PlanePowers is the only enterprise-scale knowledge and work management platform for self-sovereign AI conscious businesses — agents and teams working together to accelerate progress. 🎎🤖📈
Show more
Open weights is the new open source!
20 million ~ H100 GPUs equivalent of compute exists globally. 5% ~ of this training for 6 mo roughly produces a 10 trillion parameter SOTA model.. Astra / Fable level. The vast majority of compute is used for inference.. pre-training is billions of times less efficient than biology.
Show more
There are only 7 startups in history to cross $1 billion in ARR in < 6 years flat … @HelloSurgeAI @mercor @togethercompute @AnthropicAI @FireworksAI_HQ @wiz_io @cursor_ai … @planepowers and @liquidai will be on this list. It’s very humbling to have been the founding investor in both, from zero. The founders of these companies are truly spectacular humans.
Show more
Making on-device AI even fast + all open-weight
Today, we release DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality. A lightweight draft model proposes a block of candidate tokens and the target model verifies them in a single forward pass. Across MATH500, GSM8K, HumanEval, MBPP, and MT-Bench at batch size 1: > Up to 3.18x throughput on an H100: LFM2.5-8B-A1B on MATH500, 428 → 1362 tok/s > Up to 2.87x on an M4 Max MacBook Pro: LFM2.5-1.2B-Instruct on HumanEval, 136 → 389 tok/s > LFM2.5-2.6B means: 2.67x on the H100 (323 → 864 tok/s), 2.27x on device (61 → 139 tok/s) > Under greedy decoding, the emitted sequence is identical to baseline by construction, so benchmark accuracy is unchanged. Each draft model is around 300M parameters, with embedding and LM head tied to its target model. The gain shows most in agentic workloads, where the model reasons before every tool call and the user waits through it all: on BFCL multi-tool scenarios, DSpark cuts LFM2.5-2.6B latency by nearly 50% on average. 🧵
Show more
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
Show more
0
234
6.3K
881
Forward to community
Totally agree. Penrose OR is the fundamental process to which Schrödinger referred. OR was present in the early universe before life began. Polyaromatic hydrocarbons from stars were plentiful, coupled, oscillated and had OR moments including pleasure. Life began based on this organic quantum optics and developed as a vehicle for consciousness, as written in ancient Indian writings.
Show more
1/10 Agents and teams share the work, not just the workspace.
Godot gives studios something no other major engine does: their own technology, on their own terms. We build the commercial services around it. We're hiring across multiple roles: - Engineer (Professional Services), NA or Europe - Solutions Architect, NA - Business Developer, NA - Head of People & Culture, Europe All remote. #GodotEngine# #gamedev#
Show more
Looking for 2-3 AI companies that are already SOC 2 compliant but frustrated with their current platform + want to join the likes of @opencode, @dub, @better_auth - Dedicated team that will migrate them - 24/7/365 Slack support - Approval of @compliantvc Tag them below 👇
Show more
yesterday we made them more compressed! today we make them faster than ever with speculative decoding! up to 4x decode speed up on device for our 1.2B, 2.6B and 8B moe. You gotta try these LFMs for function calling applications on device or latency critical load on the cloud! work of art by our very own @tugot17 enjoy 🚢🚀
Show more
Today, we release DSpark draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B. These add a speculative decoding path that trades a minimal memory increase for a large decoding speedup without changing output quality. A lightweight draft model proposes a block of candidate tokens and the target model verifies them in a single forward pass. Across MATH500, GSM8K, HumanEval, MBPP, and MT-Bench at batch size 1: > Up to 3.18x throughput on an H100: LFM2.5-8B-A1B on MATH500, 428 → 1362 tok/s > Up to 2.87x on an M4 Max MacBook Pro: LFM2.5-1.2B-Instruct on HumanEval, 136 → 389 tok/s > LFM2.5-2.6B means: 2.67x on the H100 (323 → 864 tok/s), 2.27x on device (61 → 139 tok/s) > Under greedy decoding, the emitted sequence is identical to baseline by construction, so benchmark accuracy is unchanged. Each draft model is around 300M parameters, with embedding and LM head tied to its target model. The gain shows most in agentic workloads, where the model reasons before every tool call and the user waits through it all: on BFCL multi-tool scenarios, DSpark cuts LFM2.5-2.6B latency by nearly 50% on average. 🧵
Show more
2021: I start a podcast senior year of Stanford. @alexatallah is the very first guest. His frame of thinking back then rings even truer today. "Any time information is being transferred in the internet, value can be transferred too. Just imagine every website getting marketplace-ified, every community on the internet has members that want to exchange value. And it's just a question of figuring out what that value is. Now that we have a really easy way of of building that value, it's just about creativity now and figuring out what's gonna stick. And there is so much incentive to answering that question. "
Show more
So @eboyden3’s group just built the best microscope anyone has ever pointed at a vertebrate brain, and it’s pointed at the wrong thing. 200.8 Hz volumetric, 5.86 µm z-step, one number per soma every 5 ms. That’s the neuron’s action potential domain of resolution. But we know the real and fundamentally ultimate computation in the brain is happening in the cytoskeleton at kHz, MHz, GHz and even THz frequencies — six to nine orders of magnitude above anything this instrument can touch. And you don’t have to take my word for the gap, it’s in their own data: they recorded a plane at 1 kHz, downsampled to 200 Hz, and lost 30% of the spikes. That’s a third of the coarsest possible observable, gone. Then it gets worse... The indicator is Positron2-Kv — the Kv2.1 sequence is there specifically to keep it in the soma. So a plasma-membrane voltage sensor, already blind to tubulin conformational states, dipole dynamics, C-terminal tail ionic waves, tryptophan networks, is further engineered to avoid neurites, which is where the microtubule mass actually is. They then cap axial range at 170 µm and say so explicitly : the ventral brain “is dominated by neurites.” The substrate is explicitly labeled as the part not worth imaging. Whatever makes it through acquisition gets stripped by the pipeline anyway — VolPy high-passes at 15 Hz and adaptive-thresholds to spike trains, so anything subcellular or above 100 Hz is deleted by construction. And the artifact filter is the part that really perturbs me: they exclude 8–41% of ROIs for showing “highly synchronous pulse-like” activity with spatial gradients along one axis. That is exactly what a real coherent brain-wide event would look like. The rejection criterion and the hypothesis aren’t separable. Net: ~10⁷ bits/sec off a ~10¹⁴-element system, in a fish paralyzed with pancuronium, under 33× the excitation power of calcium imaging, 50% photobleached in 210 seconds. None of this is a knock on the engineering. My comments here are about what the membrane-potential paradigm can and can’t reach… nothing in existence hits sub-µs temporal, sub-100 nm spatial, cytoplasmic rather than membrane, non-perturbative, and intact-tissue at the same time.
Show more
MIT researchers have invented a new microscope that can image electrical activity in neurons distributed across the brain of an entire organism — in this case, the experimental model Danio rerio (zebrafish). Professor @eboyden3, the senior author of the study, holds a joint appointment in Media Arts and Sciences (the academic program at the Media Lab) and @mcgovernmit. The study was published in @naturemethods.
Show more
Does anyone know whether the @Etched chip will support @__tinygrad__ out of the box on Day 0 of General Availability?
50+ Fortune 500 companies run Plane in production. Plane is the most malleable, extensible, and sovereign work platform you can deploy across your enterprise — for humans and agents. Control your data. Control your intelligence. Control your future. @planepowers for the win.
Show more
The best startups are similar.
Richard Feynman said that one of the joys of working with all the really smart people on the Manhattan Project was: “You never had to explain anything twice. Everyone was paying attention. Everyone understood the first time you said it.”
Show more
Try out our new Q4_0 checkpoints for the LFM2.5 series trained with QAD. They recover a fair amount of quality lost to standard quantization techniques. 230M: 350M: 1.2B: 2.6B:
Show more