Register and share your invite link to earn from video plays and referrals.

Search results for BatchInference
BatchInference community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including BatchInference
Run inference over millions of records — free of SQL, and without your data ever leaving Snowflake. Here's distributed batch inference at scale ⚙️ Title: Batch Inference at Scale URL: ⚙️ Overview A capability that runs distributed inference workloads on Snowpark Container Services (SPCS) with Ray as the execution framework. Inference runs as a dedicated distributed workload, supporting both traditional models and LLMs, consolidating complex operations into a single API call. ❓ Challenges Solved Many customers, especially those migrating from non-SQL systems, need batch inference decoupled from SQL. ・This is especially true for files and unstructured data at large scale ・Rearchitecting workflows around SQL-first patterns is a heavy burden 💡 Methodology & How It Works ・The input DataFrame is materialized and written to a stage as Parquet files ・A job is provisioned on SPCS; the primary node initializes as the Ray head and replicas join as workers ・Each worker reads staged data, performs inference independently, and writes results to an output stage ・Unified API: a single run_batch() call handles both structured and unstructured data ・Multimodal support (images, audio, video); workers load weights once and reuse across batches; JobSpec controls workers and GPU allocation 🌍 Use Cases ・Nightly summarization of millions of support tickets ・Product catalog enrichment via image-to-text generation ・Information extraction from scanned PDFs, audio transcription and labeling, video classification and description BatchInferenceTask integrates with Snowflake Tasks for DAG automation, and all processing stays inside Snowflake — running large-scale inference while preserving data governance. #Snowflake# #BatchInference#
Show more
introducing preemptible compute for together gpu clusters same nvidia gpu infrastructure, 50% of the on-demand price built for evals, fine-tuning, batch inference + short experiments, with up to 5 minutes to checkpoint before reclamation now in public preview
Show more
The Compute Desk Nvidia B300 GPU-hour index is at its all-time high. As margins on training grow thinner, neoclouds are collectively shifting focus to inference. High-priced B300s are generating ROI by lowering the marginal cost per token generated through batch inference.
Show more
$SKHY Is Expanding Beyond HBM SK Hynix is developing Processing-in-Memory (PIM), which processes data inside or near memory to reduce data movement, latency, and power consumption The company claims PIM can deliver 300× the memory capacity of SRAM in the same chip area and is exploring hybrid bonding to stack compute and memory dies without sacrificing density Meanwhile, High Bandwidth Flash (HBF) combines 3D NAND density with HBM-style interconnects to deliver higher capacity at lower cost, targeting long-context and large-batch inference SK Hynix is also developing SALT-KV, software that distributes KV cache across HBM, DRAM, and SSDs to reduce dependence on expensive memory
Show more