Register and share your invite link to earn from video plays and referrals.

Derek Colley
@DerekColley_
CTO building with agents & open models. Automation, robotics, speech Building - -
365 Following    315 Followers
You can make Qwen 3.8 27b look smarter, dumber, faster, slower or completely broken without even changing the weights. I even locked qwen3.8 in a clean-room with a mystery executable and made it reverse engineer the program from behavior alone. QWEN 3.8 27B IS NOT ONE NUMBER. I spent the last several days absolutely ripping the model apart. 5 x DGX Sparks running almost 24/7 since it was released and the 5090s zipping through its own test suites. I still have a lot of really fun tests that im throwing at it, but i wanted to churn through the big ones first and get some of the interesting findings out. So far ive come out of it trusting single benchmark numbers a lot LESS. I went deep into the things everyone is talking about right now: thinking OFF vs LOW vs XHIGH. what “MEDIUM” actually does. why more thinking can make results worse. how reasoning can eat the answers tkn budget. why old thinking can silently balloon agent ctx. temp and determinism. BF16 vs FP8 vs NVFP4. “262K context” vs the context you actually get. speculative decoding. DSPARK. MTP. 4-bit KV cache. llama.cpp vs SGLang vs vLLM and what happens when the runtime itself is the thing breaking the model. Some of it was actually "holy shit" territory. I had the same qwen weights more than double in speed on the same 5090 by just changing the inference system around them. I had max reason lose to LOW, then found the reason wasnt simply that LOW was smarter..I had a server configured for 262k context that had nowehre near 262k of realized capacity. I had a very fast inference setup produce absolute shit until one Blackwell specific fix changed the result completely. I also threw away half a day because the harness failed , not the model. The biggest lesson from all of it: the weights are only one component of the system if you give me a model benchmark without the runtime, template, reasoning police, token budget quant and serving stack youre just giving me half a result. Full write up, graphs configs failed runs and reproducible evidence below
Show more
I circled back to @bot... and I'm FAST becoming a convert! I asked my sales bot to create an explainer based on a ppt file...
Qwen3.8-27B NVFP4 + MTP is live for GB10 🚀 Quantized with @NVIDIAAI ModelOpt 0.46.0rc1 using its shipped qwen3_5 recipe. 2.45× c1 speedup vs AR 84.3 tok/s at c8 262K context, 8/8 NIAH 17/17 zero-error runs Model: Repo:
Show more
Maya Angelou said it right: “Success is liking yourself, liking what you do, and liking how you do it.”
0
17
2.8K
435
Forward to community
I’m in love with this sentence: “Every pattern in your life repeats until you learn the lesson. The moment you choose differently, the loop ends and growth begins.”
0
209
42K
9.3K
Forward to community
A lot of people on X are new to the markets. So I'll make this as simple and harsh as possible. The economy is not going up. It is not going down. It is doing both at the same time. This is called a "K-shaped economy" and almost nobody talking about affordability on your timeline understands what it actually means. If you own assets, stocks and real estate, you are on the upper arm of the K. Your wealth has compounded aggressively since 2020. The S&P is at all-time highs. Home prices are up 50% since the pandemic. Your 401k is the best it has ever been. If you do NOT own assets, you are on the lower arm. Rent is higher. Groceries cost more. A home that cost 2.2x the median income in 1960 now costs 5 to 7x. You are working harder and falling further behind, and it is not your imagination. The top 10% of Americans own 93% of all stocks. The bottom 50% own 1%. The top 10% hold 68% of total wealth. The bottom 50% hold 2.5%. This is not an opinion. This is Federal Reserve data. When the S&P goes up 25% in a year, the people who own stocks get 25% wealthier. The people who don't own stocks get nothing. That is the K. Same economy. Two completely different outcomes based on one thing: whether you owned assets before the run started. This is why your timeline is split in half. Half the people saying the economy has never been better. The other half saying they can't afford to live. Both are telling the truth. They are just on different arms of the K. How long does this last? The honest answer is that it has been going on for decades. The Richmond Fed traced this pattern back 30 years. It happened after 2001. It happened after 2008. It happened after COVID. Each time the gap widened and never fully closed before the next shock hit. The last time wealth concentration looked like this was the Gilded Age, 1870 to 1900. That lasted 30 years before structural reform changed anything. What does this mean for you? The K-shaped economy is not a reason to panic. It is a reason to understand where you sit and what you can control. The people on the upper arm did not get lucky. They owned assets. That is the entire difference. Change your life and start small. No matter how small your account balance is. I only started with $10K, twenty years ago. This is your time to retire your whole family. They are counting on you.
Show more
0
216
4.6K
560
Forward to community
Audi CRT diesel seizes at 200,000 km. 28,000 EUR with VAT. Qashqai J12 1.3 MHEV, 2021, runs itself out of oil. 15,000 EUR with VAT. GLC 220 d 4MATIC Coupe spins a crank bearing. 26,000 EUR with VAT. A BMW i3, which runs the same electric motor across every series from 2015 to 2022, costs 909 EUR with VAT. New. Read that again, because it is the part nobody tells you when they explain how expensive electric cars are to own. We have 31 new electric motors and 11 new gearboxes on the shelf. Available today - even if it rarely fails - audi need most the help. Now let me explain why we are holding this much stock, because it is not the business we are actually in. We repair these units. That is the project. Stator rewinding, bearing work, resolver faults, the whole drive unit down to component level, plus the tooling that has to be built from scratch because nobody else is building it. That work takes time. Sometimes weeks, depending on what walks in the door and how much of it we have to reverse engineer first. If your i3 is your only car, if you run a fleet, if the car is a taxi or your work depends on it, "we will have it back to you eventually" is not an answer you can live with. So we keep new drive units in stock for exactly those people. You get the car moving now, and your old unit stays with us. That is the trade. Which brings me to the second part. We are selling the i3 eDrive system essentially at purchase price. No margin games. On top of that, if your old unit has never been opened, never been "inspected" by somebody with a hammer, completely untouched, you get 10% off. Why untouched? Because a unit somebody already cracked open is useless to us. Failed drive units are what we train our technicians and our franchise network on, and they are what we build diagnostic and repair tools from. Once a unit has been disturbed, the failure evidence is gone. It tells us nothing. So the deal is honest in both directions. You get back on the road today at a price no one else in Europe is offering. We get the failed hardware we need to keep pushing this forward. And every dead motor that lands on our bench moves us closer to the day this part gets repaired for a few hundred euro instead of replaced for a few thousand. That is the entire point. The stock on the shelf is just what we do for people who cannot wait for it. If you are not in a hurry, send us yours and let us look at it first. Plenty of what people write off is repairable. In stock. Shipping immediately. &
Show more
The median company is spending $12 / employee / month on AI The top 1% are spending $7,500 / employee / month Not sure we've ever seen an adoption gap quite like this (h/t @tryramp data, @a16z)
Show more
0
150
2.5K
296
Forward to community
London is not only dominating in many AI applications (@synthesiaIO & @ElevenLabs ) our labs are now going toe to toe with the frontier labs. @inherent_labs is one example. @Orbital_Ind is another. Then there’s labs like @Recursive_SI and @IneffableLabs which between them have raised nearly $2bn. Plus companies like @CallosumAI and @CosineAI which are crushing it as well
Show more
Cursor is now part of @SpaceX. Today, we have officially closed our acquisition. We will join the @SpaceXAI team to help make Grok the world's most useful AI and improve Grok Build, Grok Bot, Grok API, Cursor, and more. SpaceX has built some of the most inspiring and impressive technology in the world, and we’re grateful for the opportunity to become part of such a special company. Onwards.
Show more
0
2.4K
55K
4.6K
Forward to community
if you're on llama.cpp run Qwen3.8-27b with -spec-default --spec-type draft-mtp Since it does take its sweet time to think, might as well let it think fast. Run with mtp, you don't need a sep drafter, getting 2x speed on decode now, totally worth it. full command i'm using on my 3090 ./build/bin/llama-server -m "/qwen-3.8/Qwen3.8-27B-Q4_K_M.gguf" --host 127.0.0.1 --port 8080 -ngl 999 -fa on --jinja -np 1 -t 12 --alias qwen3.8-27b-q4 --spec-default --spec-type draft-mtp --cache-type-k q8_0 --cache-type-v q8_0
Show more
find & replace, but for a voice recording 🎙️ FireRedTTS3 is out on @huggingface, an omni-TTS model that can do voice design, voice cloning and a cool new feature: speech in-painting - or voice editing! ▶️ on Spaces
Show more
Qwen3.8-27B can now be run locally! ✨ Run on 17GB RAM via Unsloth Dynamic GGUFs. Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. GGUF: Guide:
Show more
0
202
5.2K
620
Forward to community
I just watched "3 body problem" (netflix). Spoiler alert - look away if you want to watch it... ______________________________ The San-Ti (aliens) deploy 2 pairs of Sophons (quantum entangled) to monitor humans. Anything a sophon experiences, instant replays 400 light years away... instant communication with a superior being. Leaving aside the alien angle, this (quantum) is how we could communicate across the galaxy🤔
Show more
🦔Sam Altman has been describing a version of ChatGPT that remembers your whole life, every email, conversation, and document you choose to feed it, so it can act as a proactive assistant. A viral clip framed this as AI surveilling your entire life within six months, which overstates both the timeline and what he said. Strip out the hype, though, and there's a business angle here, one Altman has said out loud. My Take The obvious worry is privacy, and it's a big one. You'd be handing a single company a record of your entire life, your emails, your health, your money, your private conversations, and trusting it to guard all of that and never misuse it. These are the same firms that keep getting hacked and subpoenaed, and once that data exists in one place, you don't control where it ends up. They want it anyway because it locks you in. Altman has said the deep memory creates massive switching costs, which is his term for a moat. Once ChatGPT holds years of your context, walking away for a competitor means starting from zero, so most people won't. You carry the privacy risk so they can hold a customer who won't walk. Google and Apple ran a version of this with email and photos, and this hook goes far deeper, your whole life instead of a photo library. I'd think hard before pouring all of it into any single one of these tools, because the company that holds your entire context doesn't have to stay the best. It only has to be impossible to quit. Hedgie🤗
Show more
Why would anyone pay SpaceX or Nebius $30-$50B/year/1GW of compute? Because OpenAI and Anthropic can generate $100B+ per gigawatt per year selling inference API! Here's how that's possible, and why it's sustainable (save this) First, what are they actually selling? Every ChatGPT answer, every Cursor autocomplete, every enterprise copilot runs on "inference", the model generating tokens. The labs sell those tokens through subscriptions or an API, metered like electricity @SemiAnalysis showed their model: run one gigawatt of Nvidia GB300 compute selling tokens at posted API prices and it generates over $100 BILLION a year of revenue. That same gigawatt costs roughly $12B-$50B/year to rent out A 2-8x spread between what compute costs and what intelligence sells for So why can they charge that much? Because the customer isn't comparing token prices to compute prices. They're comparing tokens to LABOR A few dollars of tokens replaces work that costs hundreds of dollars an hour. Legal review, sales ops, financial close, code. That's why enterprise agent adoption is up 20x to 108x across job functions in just five months. At today's prices the buyer's ROI already fantastic Why it's sustainable: 1. Demand compounds faster than prices fall. Token prices drop constantly, but agents burn dramatically more tokens per task than chatbots ever did, and every job function is adopting at once. Falling price x exploding volume = growing revenue 2. Supply is rationed. A handful of frontier labs, and none of them have enough compute. When you're capacity constrained you serve the highest-value demand first and pricing holds 3. The buyers keep paying UP, not down. Microsoft sells this same inference through Azure and Copilot. Nebius just disclosed its first deal at $40-50M per megawatt, the top of its own range. Nobody negotiates prices higher on a product that's about to be oversupplied This is why the "AI capex bubble" framing keeps missing. The $100B at the top of the stack is what pays the $30-50B compute deals, which pay the datacenters, the chips, the memory, the power. The most profitable product in tech is funding everything below it And you don't need to own the private labs to win. Every dollar of inference revenue flows down through the infra stack, and that's exactly where I'm positioned (compute, memory, power) If this was helpful, my company provides a service where 5 top-tier analysts share their market analysis and real-time portfolios so you can see exactly how we're positioned across this stack. It's inside Milk Road PRO and just $1 to try (insane price just to check it out). Learn more here: Follow me @kylereidhead for more insights on AI, robotics and markets!
Show more
Every serious TTS ships emotion tags like [whispering] [laugh] [excited]. Almost nobody uses them - nobody wants to prompt-wrangle prosody. Expressive mode fixes that: the LLM writes emotion inline, Fish Audio S2.1 Pro renders it, one toggle in @livekit Agents. Congrats to the LiveKit team on the launch! And proud of the collaboration that made this sing. 🐟
Show more
Sarvam AI opened Voice Agents today, so I built a Hyderabadi sabzi aunty 😭 Called her and bargained in Hindi + Telugu + English, kept interrupting, switched languages mid-sentence, and she still remembered my order, gave me a final total, and even confirmed my (fake) UPI payment. Took me ~15 mins to build. It was fun to recreate the lost art of vegetable bargaining, but this time with an AI agent.
Show more
0
56
1.5K
114
Forward to community
Got some merch... At the TS AI conference in London @mastra