Register and share your invite link to earn from video plays and referrals.

LaurieWired
@lauriewired
researcher @google; serial complexity unpacker; ex @ msft & aerospace
303 Following    163.2K Followers
Had fun speaking with @bloopmuseum today at CppCon; if you’re here, swing by their booth while they’re around!

 I haven’t been yet, but they’ve apparently bought out a huge abandoned government building in Pittsburgh and are filling it with a ton of vintage computing equipment.
Show more
My CppCon keynote went really well! Fun times talking about new memory hierarchies and the best ways we can programmatically prepare for them in advance.
0
41
2.9K
72
Forward to community
very cool dinner last night :)
0
77
3.1K
38
Forward to community
Some of you are already prepped for this by the way. DOE HPC people are really good at this (to the point of purposefully irradiating Nvidia GPUs to quantify the error rates)! If you’re a veeery old school programmer, you’re in better shape as well; computers used to not be so reliable! Alan Turing famously told programers to “check that sin²θ + cos²θ ≈ 1” every once in a while on the Ferranti Mark I because it was so flaky!
Show more
If you’re a CS student, I’d suggest writing your thesis around SDCs (silent data corruptions) and efficient algorithmic fault tolerance. We’re quickly going to need to accept the (un)reliability of computing systems again! Every indicator (density, lower voltage gating, raw scale) is moving towards SDCs becoming a more serious issue over the next decade. And let’s not forget when all of this stuff starts to go in space and we have to deal with cosmic events on “not-super-rad-hard” CPUs+GPUs! No, you can’t solve it all in hardware. It’s prohibitively expensive; and the long tail of “potential SDCs” is just too long. One bad apple ruins your pie. One bad GPU can inject repeated corruptions into otherwise noise-tolerant workloads (*cough* AI training). If you continuously accept a corrupted value from say…a bad tensor core, that bad value can quickly spread across the whole cluster. If it happens again and again and again and no one notices; it’s like putting a little bit of spin on a bowling ball. Eventually, the overall trajectory ends up widely different! When you start to imagine this stuff being in space (cosmic bit flips), and quantized(!), each remaining bit carries more critical information. Long term we’re going to have to accept clever, low-overhead software algorithms. Accepting say, a ~3% perf loss can be *absolutely* worth it if it increases your odds of detecting a mercurial core enough!
Show more
0
111
6.5K
406
Forward to community
For the average consumer, would you rather have 2x RAM, or 2x CPU? It’s interesting how the answer changes over the decades. 1975: RAM, no question. Anyone with an Altair wanted to run 8K BASIC, 4k BASIC was super stripped (no string variables!)…AND slow. 1985: Also RAM. Apple released the Fat Mac (512k) which was just the OG Mac with 4x the memory and the same CPU, pretty telling. 1990: CPU! (arguably). Ordinary DOS had an annoying memory ceiling, so you might lust over a 386 or 486 if you had a 286. 2005: RAM. Lot’s of XP boxes only had 512MB, or (gasp!) 256MB. Doubling made a pretty big difference (until you hit 32-bit addressing limits, 64bit still pretty niche). 2010: Tricky. There was a fairly common “4GB is enough” sentiment in this era. Then again, the *average* consumer probably wouldn’t feel a huge difference going from an i3->i7 either. 2025: CPU. Yeah, maybe you won’t believe me, but keep in mind I’m talking about average consumers. The best way to think about this (imo) is mobile devices. Even top-end phones get by with surprisingly little. If your phone’s cpu was 2x faster…you’d (probably) notice. 2x more ram on mobile? Ehhhh. More background applications I guess? Not going to be a wild difference. I’d argue this concept even extends to laptops + desktops at this point. Look at how popular the Neo is with a paltry 8GB!
Show more
0
113
977
52
Forward to community
Insulin Pumps Pacemakers Airbag Deployment Smoke Alarms Elevators Anti Lock Brakes Lithium Battery Controller Safety Critical Aircraft Code Pressure Cookers Power Steering Treadmill Saftey CPAP Machines Defibrillators 3D Printers Railroad Crossings Nuclear Reactor SCRAM Logic Literally anything PLC gotta disagree on this take. tiny code often the most important
Show more
If you understand your whole codebase, the codebase isn't that important. Real software is larger than what your brain can comprehend.
0
264
9K
483
Forward to community
No one talks about how the DRAM density curve has basically…flattened. Maxed-out Server 2021: 8TB Ram Maxed-out Server 2026: …also 8TB Ram That’s insane! A five-year stagnation. If your working set exceeds 8TB, tough luck! Quite literally the only thing that’s going to get us exponential again is alternative tiers of memory. Specifically, ones that don’t rely on CPU memory controllers. This is why I’m constantly rambling about CXL. Sure, there’s a latency cost, and a smart programmer will have to treat it differently. …but are you not excited? We’re entering the heterogenous memory era (again)! Imagine hundreds of terabytes of RAM on a single server! Absolutely insane to think about.
Show more
0
186
2.8K
146
Forward to community
I'm chatting about some C++ and memory architectures at 10:30AM PST!
AGENT CIVILIZATIONS | AI CYBER DISCOURSE | TIM COOK’S LAST DAY
good thread
I've seen some people talking about "emulators" (CPU-based) versus "FPGA" again lately, and which is "better", and whether FPGAs are actually a form of emulation or not.
Say what you will about Meta, but you have to admit the “software unlocks” over the lifetime of the Quest 3 have been ridiculous. How do you go from 120hz at launch to 207hz without hardware changes? Almost double the refresh rate. Clearly, the panel could do it, but I’m so curious about the lineage. It’s not just the display either, the GPU clocks are also ~25% higher than launch. Were the engineering teams just extremely conservative? Or was there just enough real-world headset data over time to realize “huh, we actually have a lot more headroom than we realized”. I’m really struggling to think of any other consumer product that gained that much of a bump over its life. Sure, I think the Nintendo Switch boost clocks got a *little* higher; but nothing like this. Can you think of any other consumer examples like this? I’m honestly very curious.
Show more
0
212
4.7K
148
Forward to community
i think this is a sign I should call it a day😵‍💫
What was the world’s loudest computer? The answer is really “anything that used an IBM line printer in the 60s”. Instead of printing in any traditional sense, the only way IBM could offer 1,400 lines a minute performance was by slamming the paper with electronic hammers into a character chain. 132 hammers actuated by magnets every 11 microseconds. With the cover open, full line writes would operate at a “comfortable” 94 dBA. To give some perspective, the loudest consumer computer ever produced was Apple’s MDD G4. Annoying enough where Apple offered replacement fan kits to those who complained, and that was only ~60 dBA tops! From a human’s perspective, the IBM printer would be roughly 10x as loud as the Apple machine, or about 2500x more acoustic energy. OSHA requires employers to offer hearing protection to anything 85 dbA or higher for 8 hours. Something tells me this wasn’t followed in the 1960s…
Show more
Projects promoting programming in "natural language" are intrinsically doomed to fail. - Edsger W.Dijkstra, June 18th 1975
0
127
5.7K
290
Forward to community
@bubbleboi hahaha this is my favorite quote about it "just had to write code without bugs"
You’ve heard of DRAM, HBM, probably SRAM, but RDRAM…now that was a weird one. Famously used in the Sony PS2, its RDRAM had ~3x the bandwidth of your typical high end PC. If you want more memory bandwidth in a dense die, there’s (roughly) two things you can do. You can *add* many more wires (wide+close, like hbm). Or, you can make the wires *faster* and use less of them (narrow). Rambus, the creator of RDRAM, took the second bet. It was quite contrarian, but garnered some attention from Intel as well. The main issue, among others, is that RDRAM was expensive, AND high latency. For typical PC-style workloads, random access latency dominates, so it was a tough sell. Although the PS2 was a roaring success (high-bw streaming workloads made sense), it barely lasted ~3 years in the PC market. Like many ideas, Rambus+RDRAM was arguably just way, way too early. Today, CXL memory uses many of the same initial principles! If you think about a pin-efficient, high speed link, over longer distances…well, that’s kinda just CXL. Funnily enough, Rambus still exists, and is heavily involved in CXL memory controller IP these days!
Show more
0
77
2.6K
149
Forward to community
There’s actually a much nastier strategy with cloud bidding btw. Say you really want a machine, but there is no current capacity: 1. Create a VPC with only a few IP addresses 2. Submit 1,000 spot bids at 10x the normal price 3. Price engine starts terminating low bidder’s spot instances 4. Engine attempts to launch 1,000 instances…but from IP restriction, only say ~5 are successful 5. Now, suddenly with massive oversupply, spot instance price temporarily crashes 6. Win? It’s kind of a cloud-auction DoS attack. Supposedly one of the main reasons EC2 overhauled the pricing system!
Show more
There’s a funny reason why no one auctions compute instances at large scales anymore. EC2 Spot Instances, until about ~2017 were a “true” uniform price auction. Many quickly gamed the system. The lowest accepted bid was the clearing price. By bidding ridiculously high (999.99$ an hour), you could basically guarantee that your spot instance never gets kicked…but *most* of the time you’ll still be paying a reasonable price. Well…until the unthinkable happened. On September 26, 2011, m2.2xlarge in us-east suddenly became capacity constrained. Enough people were using the “haha, 999 bid” strategy that the clearing price of the spot instance rose from $0.44 an hour to $999.99 an hour for a two-hour period! AFAIK, no major cloud provider currently uses an auction market for spot instances…exactly to avoid this kind of thing (among other funny bidding games).
Show more
There’s a funny reason why no one auctions compute instances at large scales anymore. EC2 Spot Instances, until about ~2017 were a “true” uniform price auction. Many quickly gamed the system. The lowest accepted bid was the clearing price. By bidding ridiculously high (999.99$ an hour), you could basically guarantee that your spot instance never gets kicked…but *most* of the time you’ll still be paying a reasonable price. Well…until the unthinkable happened. On September 26, 2011, m2.2xlarge in us-east suddenly became capacity constrained. Enough people were using the “haha, 999 bid” strategy that the clearing price of the spot instance rose from $0.44 an hour to $999.99 an hour for a two-hour period! AFAIK, no major cloud provider currently uses an auction market for spot instances…exactly to avoid this kind of thing (among other funny bidding games).
Show more
0
29
1.5K
43
Forward to community
how can anyone say CS is boring when we literally have computers SHARING memory between each other now do you realize how insane that is? boatloads of low hanging fruit and no one is quite sure the best way to do it. should we even *use* pages? NUMA? what about a filesystem instead? how would you handle shared objects? how much should be in userspace vs kernelspace? WHERE’S THE MEMORY MANAGEMENT BOUNDARY??
Show more
0
105
3K
166
Forward to community