Register and share your invite link to earn from video plays and referrals.

Search results for Optimization
Optimization community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Optimization
Harness optimization is getting real receipts. AutoSaddler is offline harness learning from agent failure traces. Not another prompt tweak loop. It patches prompts, tools, and middleware as code. Then it keeps updates that survive a held out set. On the test sets (Pass@1): GAIA2: 53.0 → 62.0 (+9.0) SWE-Bench Pro: 37.3 → 46.9 (+9.6) Terminal-Bench 2.0: 40.0 → 50.0 (+10.0) That TB2 number also clears the expert tuned Terminus KIRA at 47.5. Kill generalization aware selection and GAIA2 falls to 50.6, under the default agent. Deep diagnosis and structured patches help. Dev set filtering is what stops the harness from overfitting the mini batch. On GAIA2, Figure 1b, about 147 leveraged traces to the best dev score vs about 1,400 for Meta-Harness.
Show more
Another optimization has landed in Plonky3: proof size reduction through Merkle path pruning, now for FRI. Feedback and suggestions are welcome 👇
5 FABLE 5.1 OPTIMIZATIONS YOU CAN USE RIGHT NOW: 1. Set effort to low 2. Run /claude-api cost-optimize 3. Run claude-api prompt-audit 4. Change effort mid-conversation without a cache hit 5. Update your API config with claude-api migrate
Show more
Interesting paper on prompt optimization. They claim that a single-lineage prompt optimizer just matched GEPA on a smaller rollout budget. Prompt optimization has been drifting toward heavier machinery, with candidate pools, reflection trees, and Pareto-based selection. NPO keeps one lineage. At each iteration it runs the student on the current prompt, collects rollout traces and rewards, and hands a sliding window of recent iterations to a teacher model that rewrites the prompt. There is no candidate population and no search tree. On the two instruction-following benchmarks it spends 3,500 and 6,800 rollouts against GEPA's 3,593 and 6,871, and it stays broadly comparable across 22 TextArena games. The interaction with teacher strength is what makes this interesting. NPO's advantage grows as the teacher model gets stronger, which suggests optimizer-side search complexity has been compensating for weak teacher reasoning all along. Paper: Chat with Paper:
Show more
have the new codex optimizations improved your usage?
Has the optimization backlash begun? @amandamull joins Bloomberg This Weekend to talk about why people are embracing imperfection, saying social videos can “bring everybody this flattened average of a perfect face and I think we’re sick of that.”
Show more
Ass-covering optimization 🤣 “There’s a tremendous bias against taking risks. Everyone is trying to optimize their ass-covering” -Elon Musk
0
170
657
103
Forward to community
Our Single-rollout Asynchronous Optimization (SAO), is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench.
Show more
0
50
1.1K
109
Forward to community
$NVDA says Blackwell software optimizations improved DeepSeek V4 performance by ~5x in one month cutting token costs to ~20% of prior levels. Nvidia’s value keeps compounding after deployment as software, CUDA, networking and hardware drive more throughput from the same GPUs.
Show more
Nature's optimization algorithms are more ruthless than investors Sometimes I think they're just better at finding value