Register and share your invite link to earn from video plays and referrals.

Search results for CUDA
CUDA community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including CUDA
From CUDA to ChatGPT, $NVDA’s rise has mirrored every major wave in modern computing. Now, optimism around US-China AI chip talks is pushing the company to a new record market cap of $5.6T. Can the AI boom keep powering the next leg higher for $NVDA?
Show more
Can AMD break the CUDA Moat? AMD Advancing AI 2026, Up to 105% Equity Rebate Discounts for OpenAI, Agentic Kernel Generation, Improvement in Software Quality, Unstable Internal Development Clusters, Helios MI455X Production Ramp Hell
Show more
Scaling hybrid workloads with NVIDIA CUDA-Q is enabling teams across the quantum ecosystem to explore real-world impact using GPU-accelerated simulation. 🚀 ⚡ Aegiq and @Quantum_Motion are advancing quantum chemistry workflows with CUDA-Q. ⚡ Classiq is drawing on CUDA-Q to explore new quantum applications in finance. ⚡ FirstQFM has demonstrated quantum foundation models on the Leonardo supercomputer. ⚡ Eclipse Qrisp (pioneered by @Fraunhofer FOKUS) and @Qilimanjaro are powering their work with CUDA-Q. ⚡ qBraid now serves as a CUDA-Q target, expanding access to a broad set of QPU providers. ⚡ QCentroid is building QuantumOps workflows on CUDA-Q, driving more efficient applications development. ⚡ Welinq is pairing its distributed quantum compiler with CUDA-Q for GPU-accelerated circuit verification. The path to useful quantum computing is hybrid—powered by accelerated simulation. Learn more: #ISC26#
Show more
Venus v0.2.4 is here. Every release pushes the GPU prover further. Six CUDA and data-movement optimizations deliver another 4.5–4.6% performance gain over v0.2.3. That brings Venus to 1.19 ~ 1.22× the performance of the v0.1.6 baseline on Ethereum workloads. Unstoppable.
Show more
🚀 vLLM-Omni v0.20.0 is out — aligned with upstream vLLM v0.20.0 (CUDA 13.0 · PyTorch 2.11 · Transformers 5.x). ⚡ Qwen3-Omni throughput +72% on H20, 32 conc (0.241 → 0.414 req/s) via talker / code2wav multi-replica scaling 🎙️ TTS faster & leaner: VoxCPM2 RTF 0.946 → 0.106 · Fish Speech Fast AR latency -53% · Qwen3-TTS / Voxtral-TTS Code2Wav saves ~3.2 GiB 🎨 Diffusion dynamic step-level batching: +7.8% throughput / -5.8% latency 🆕 New / improved: HunyuanImage-3.0, ERNIE T2I, AudioX, Wan2.2-S2V, LTX-2.3, FastGen Wan 2.1 📱 Wan2.2 on NPU production-ready: MindIE-SD, fused ops, VAE BF16, HSDP/USP — +50–60% perf 🧮 Quant expanded: Qwen Omni W4A16, OmniGen2 FP8, Z-Image FP8, HunyuanImage3 NPU, GLM-Image 🧩 Multi-backend updates across CUDA / ROCm / MUSA / NPU / XPU Check it out →
Show more
Today’s AI takeoff stands on decades of open research and open infrastructure. The Transformer. Backpropagation. ImageNet. PyTorch. TensorFlow. JAX. CUDA. Linux. Books, the open Internet, and millions of open-source codebases on GitHub that became training data. Without all these, today’s AI labs and companies simply would not exist.
Show more
$NVDA says Blackwell software optimizations improved DeepSeek V4 performance by ~5x in one month cutting token costs to ~20% of prior levels. Nvidia’s value keeps compounding after deployment as software, CUDA, networking and hardware drive more throughput from the same GPUs.
Show more
At Flink Forward Asia Shenzhen 2026, NVIDIA’s Chuan Chen shared how NVIDIA and Alibaba Cloud accelerate multimodal data stream processing for Apache Flink: “NVIDIA and Alibaba Cloud's team technically collaborate to enable the CUDA library-accelerated multimodal data stream processing of Apache Flink.” This open-source collaboration enables end-to-end, high-performance multimodal streaming architectures for AI commentary, live image-text feeds, and interactive Q&A. #NVIDIA# #AlibabaCloud# #ApacheFlink# #DataAI# #AI# #Multimodal# #RealTimeStreaming#
Show more