Register and share your invite link to earn from video plays and referrals.

Roy
@usr_bin_roygbiv
Professional anon @ mum's basement
914 Following    8.2K Followers
I'm having 4 separate conversations with people today about rtx 6000 price increases so I wanted to make a post explaining my original reasoning and general ideas on pricing for 6000s in particular. I believe max q and workstation rtx pro 6000 blackwells will be the SINGLE LARGEST appreciating asset period the next two years. I have many reasons for this coming from years of buying used pc parts and laptops as well as following previous cycles personally, and using a lot of different inference hardware myself the last year or so. - They are the single most powerful gpu for inference aggregate tok/s you can run on north american 120v besides the dgx station which uses a whole circuit. You can run 3-4 on a 20A circuit. - They have triple the vram of 5090s with a binned version of the same chip, and the highest memory bandwidth for matmuls - nvfp4 and 4 bit precision tensor cores means highest flops of any silicon generally available on 6000 and 5090 - You can slap them in any existing machine with an open pcie slot Why 6000s over 5090s then? Supply is way way lower, far fewer were manufactured somewhere on the order of 10-100x, they also have triple the vram on the same wattage. This means many models aren't runnable locally without them, and that unlock requires 4-8 of them for 750B models. 4 minimum for stuff like dsv4.1. New huawei chips china is targeting models for are 144-288gb, essentially reverse engineered h200s. Which is 3 6000s and the 6000s will be 2x faster than whats available in China while being more power efficient. Why 6000s over m5 ultra studios? Cuda and the tensor cores, also designated memory (not shared with os) and nvfp4 means higher speed and precision even if its the same memory on paper for a spec sheet. Even with perfect mlx kernels and if it was the same bandwidth (it's 50% higher than max spec m5 ultra) the gpus will never be the same speed for flops needed for inference because of the way the ALU is laid out on 5090/6000s. Why 6000s over h200/b200/b300 or dgx station? Nvidia is still manufacturing b300s, enterprise demand is all going to b300s and rubin which is what nvidia will focus 100% of their production on for likely the next 2 years minimum due to demand and economies of scale. Consumers are a tiny fraction of the market its not worth their time to serve. DGX stations will be tied to the price of ripping a b300 out of a node and rigging it up which is already cheaper when you look at the price of an 8x node compared to a dgx station. Why not AMD? CUDA and ALU layouts. Same as MLX but ROCm is in even worse shape. Even with perfect kernels on both compared with the same model and vram/bandwidth it won't be as fast for aggregate tps. Why not 5000 or 5500s? Lower memory bandwidth than 5090s and 6000s. Due to all of these unique comparisons and attributes, combined with the extremely limited supply, nvidia is not manufacturing more of them, has no incentive to manufacture more of them, while demand will continue to increase, and the general market of buyers able to utilize them, of which many have no other option for many models such as dsv4.1 to even run them locally (maybe h200s with fans or h170x which are also limited supply, but again significantly lower tps (1/2 most likely). They are 2x the speed for most of the options and have the perfect tax on second hand markets like ebay on top of the limited supply and the power efficiency and speed for actual real world usecases. I had originally modeled these to go to $32-40k by next year based on potential users and supply in north america, but I think this will happen far sooner now as supply is out at retailers in the next few weeks and no more are likely to be manufactured ever. I do not believe it is unreasonable to see them go to auction on ebay etc for $100-$120k late next year-early 2028 with current demand curves given the limited supply. This would be a 15x from their price in april. Because of the relative *perceived* niche of the topend prosumer market for manufacturing planning, contrasted against the real world market allowing basically all SMBs to run models a few weeks behind frontier locally with them on regular 120v power on any machine with a pcie slot at 2x the speed and half the power draw of any other option, (besides maybe m5 ultra which is still 33% slower) these have truly unique characteristics compared to any other card. Extremely high utility, extremely high TAM, and extremely high potential demand for what they're actually able to run twice as fast as any alternative for far more people. At the same time as all of this, they have the lowest or tied for lowest supply of any blackwell gpu. Do with this information what you want.
Show more
to illustrate the opportunity that still lies in prediction heads and how unoptimized kernels are this ~$120k machine performs the same as two six thousands that until a month ago were $20k ($30k now)
People on twitter posting 6 hours a day while at work then wonder how people could have multiple jobs
@_imdawon If you follow elon for long enough you realize his tweets about product are actually demands for his team to have ready the next morning
put yourself in a situation where you have enough work and responsibilities it's literally impossible to overthink and you're being compensated for it