Venice is going vertical. We're standing up B300 GPU clusters in our own data centers and we need a Venetian who lives at the intersection of CUDA kernels, quantization math, and production inference at scale.
If you've squeezed
@vllm_project or
@sgl_project past their limits, profiled with Nsight, shipped tensor-parallel multi-node configs, and want to do it all in service of privacy-first AI infrastructure for millions of users, this is your call.
Come build the Port City of AI with us.