Register and share your invite link to earn from video plays and referrals.

Jordan Nanos
@JordanNanos
Member of Technical Staff @SemiAnalysis_
901 Following    4.8K Followers
Prediction for hot chips: every startup and custom silicon program says something along the lines of “we gave up on trying to build a universal compiler and are just having AI write kernels for us by targeting the hardware instead of a bunch of abstractions… and get this, it’s actually working”
Show more
the npu inference repo in here confirms that they are using Ascend 910B
love when @OpenAI is open. incredible contribution to @OpenComputePrj 1. multiplane design for 800GbE. 2 tiers for 131k endpoints 2. custom MRC protocol that extends RoCE to support packet spraying 3. switch from dynamic to static routing with SRv6
Show more
cool idea from DeepSeek in their DualPath paper! instead of loading all KV's directly onto GPUs from local NVMe (or DRAM) and bottlenecking on the local PCIe bus, they can stage the KV's in the DRAM on the decode GPU servers, and then transfer the KV's to the prefill GPUs via GDRDMA
Show more