Register and share your invite link to earn from video plays and referrals.

Search results for Benchmark
Benchmark community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Benchmark
Benchmark performance does not guarantee business value. Organizations should evaluate LLMs within actual workflows such as measuring customer experience, employee well-being, productivity, and skill development. (2/6)
Show more
Benchmark results for Bonsai Q2_0 and Q1_0 on my GX10, using strict EvalPlus: Ternary Bonsai Q2_0 • HumanEval+: 152/164 92.68% • MBPP+: 305/378 80.69% • Total Plus solved: 457/542 Bonsai Q1_0 • HumanEval+: 149/164 90.85% • MBPP+: 282/378 74.60% • Total Plus solved: 431/542 Both completed all 542 generations with zero errors. Ternary recovered 26 tasks over Q1_0. Of those, 23 came from MBPP+ and only 3 from HumanEval+. For context, PrismML reports Qwen3.6-27B FP16 at 95.12% HumanEval+, 83.33% MBPP+, and 88.74% across its coding category including LiveCodeBench. That's a very very solid model for such weights and considering its file size.
Show more
Benchmarking @NVIDIAAI's Nemotron Puzzle 75B locally on the GX10. NVFP4 via vLLM's OpenAI API, MTP speculative decoding, forced 1,500-token generations. 🏃‍♀️22.75 tok/s in a single session into 88.85 cumulative at 7 sessions. 🧍Baseline without MTP: ~16.3. Scripts + setup:
Show more
Benchmark says SEC's NMS proposal is the 'most consequential' US crypto rule this year
On benchmarking long-context agentic instruction following. Agent benchmarks mostly reward reaching the answer. This new benchmark measures whether the agent reached it the permitted way, which is the question enterprise deployments care about. If you ship skills files, policy documents, or long system prompts, you have been trusting that they actually bind agent behavior. But how are you measuring all of this? Surge AI built a benchmark to actually check this. HANDBOOK.md places a standard operating procedure of 20 to 124 pages in context and grades whether it governed every action across an extended tool-use horizon. 65 tasks, five domains, ten fictional companies. Each task runs in a self-contained company environment with a file workspace plus mock email, chat, calendar, issue-tracking, and commerce services exposed over MCP. Every task mutates one of ten base handbooks, altering the specific rules and thresholds that grading turns on, so memorization does not help. Grading is fully deterministic and two-sided. 824 programmatic criteria check that required actions occurred and that prohibited actions did not. Paper: Learn to build effective AI agents in our academy:
Show more
Overnight benchmark run on GX10 for Poolside Laguna S 2.1 (Q4_K_M): - 20.4 tok/s at 1K context - 19.4 tok/s at 8K context - 11.3 tok/s at 128K context Local benchmark: - 64% coding - 60% debugging - 70% repo edits - 30% JSON - 30% review Optimized NVFP4 + vLLM/DFlash runs can produce higher throughput, especially with long generations and concurrent requests.
Show more
Model benchmark comparisons remind me of blockchain TPS comparisons in 2020
Most benchmarks test in controlled conditions. Real workloads don’t. → Concurrency → Latency → Scale → Cost under load These are what actually matter. This breaks down how to evaluate performance under real conditions:
Show more