Big news:
@dMatrix_AI is partnering with
@nvidia.
Our next-gen Raptor XPUs with our path-breaking 3D DRAM stacked memory will plug directly into NVIDIA MGX rack-scale systems via NVIDIA NVLink Fusion — delivering ultra-low latency inference for premium token services.
Inference is an infinite opportunity, and it will demand a diversity of compute. Customers will be able to run a Raptor rack standalone, or alongside NVIDIA GPUs — matching the best compute to each phase of the workload.
Grateful to
@JensenHuang and the NVIDIA team for building an open ecosystem that makes this possible. Let's go build the future!
🔗