Inkling running on a DGX B200 node with vLLM across 8 GPUs. Text, images, and reasoning from a single container.
Red Hat AI FP8 checkpoints coming soon.
Shoutout to the
@vllm_project community for getting this up and running and
@_soyr_ for the quick video.