๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ash Hart
@ashxhart
Doing my bit to bring local AI to everyone |
๊ฐ€์ž… October 2014
315 ํŒ”๋กœ์ž‰ ์ค‘    2.8K ํŒฌ
MCDMA | Metal CUDA Direct Memory Access ๐Ÿš€ If you have a Spark and an Apple Silicon Mac, MCDMA gives you a direct RDMA path between CUDA memory and Metal-side unified memory over USB-C. Registered memory, rkeys, one-sided READ/WRITE, two-sided SEND/RECV with credit flow control. Same verbs both ways, no master/slave. The Mac writes straight into CUDA-mapped memory on the Spark, and the Spark writes straight back into Mac memory. My setup takes it a little further: Spark 1 โ‡„ CX7 โ‡„ Spark 2 (prompt processing) Spark 1 โ‡„ USB-C โ‡„ Mac Studio (Decode) Spark 2 โ‡„ USB-C โ‡„ Mac Studio (Decode) Two independent MCDMA USBC links, so the Studio isn't stuck behind one cable; both Sparks move data concurrently, and it writes results back into either. Measured, every byte delivery verified: โ€ข 939 MB/s single link โ€ข 1.80 GB/s Mac โ†’ both Sparks, concurrent โ€ข 1.25 GB/s both Sparks โ†’ Mac, concurrent โ€ข 24 ยตs round-trip, 41k msg/s small-message One Spark + One Mac works. Two is just how I'm using it: DeepSeek prompt processing across the Sparks, decode on the Studio. Benchmarks, tests, Open Source, and write-up this week. @NVIDIARTXSpark @NVIDIAAI @NaderLikeLadder @msharmavikram Thereโ€™s still a lot of performance headroom here. If the currently locked USB4 controller can be allowed to train at full capability, Iโ€™d love to test how far we can push this. Please check your DMs.
๋” ๋ณด๊ธฐ