Register and share your invite link to earn from video plays and referrals.

Search results for 344
344 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including 344
Not sure what model Space Bunny is, but I am impressed. I had a chance to test it this week. It's free in OpenCode for the next week, with a 1M token context window and image input. I focused my tests on visual work, mostly starting from a single image or sketch. Here are my results. Overall, its visual understanding is excellent, and its design taste is strong. One behavior that stood out is that it checks its own work. When it has a browser, it opens the page, looks at screenshots, and fixes what it sees before it says it's done. These verification capabilities matter in long-horizon agents. From a rough wireframe with seven handwritten notes, it built a full landing page and followed every note. From one photo of a moka pot, it built a 3D model that matches the original down to the eight-sided body, the brass valve, and the camping stove underneath. It built a 3D anatomy of a mirrorless camera with 67 labeled parts, an 11-element lens, a 9-blade iris, and an exploded view. Very excited about these results. I asked it for a 60-second animated proof of why a circle's area is πr². It built the whole thing, with slices that rearrange into a rectangle, a live slider, captions, and narration. Not perfect, but still impressive. I gave it a rough hand-drawn sketch of a game level. It built a playable 3D game in Three.js that follows the sketch and all ten rules I wrote on it. Then it playtested its own jumps in Chrome and tuned the physics until every lava stone was reachable. Based on this, I would reach for it for image-to-code work, interactive prototypes, and creative coding in 3D. It’s also quite fast, which makes it great for rapid prototyping. Space Bunny generated everything in the video in OpenCode. I wrote the prompts and supplied the images.
Show more
Nice paper showing how to re-evaluate a production agent at a fraction of the cost. 200 questions, 38.5% of the full benchmark, reproduce the full score to within 1.03 points. The authors studied an analytics agent that serves tens of thousands of monthly users, using 574 historical benchmark runs split by date into calibration and held-out periods. They compared random sampling, cached results, fixed representative subsets and adaptive testing based on item response theory. Multidimensional 2PL adaptive testing gave the best fidelity. The team deployed difficulty-stratified fixed subsets instead because they are simpler to run. Those subsets transferred to five other agent families without recalibration and stayed stable with calibration windows as short as one day. Paper: Chat with Paper:
Show more
Don't sleep on using Jev-as-a-Judge for agent evaluation. This is one of the most impressive Jev use cases I have found so far. Jev is a natural fit as a Judge, but it doesn't mean you use it everywhere. Similarly, you shouldn't use frontier models for evals everywhere. I'm running lots of tests on this atm, but early results point to an optimized flow (balancing accuracy and cost) that combines Jev and frontier models. Concretely, use Jev in high-confidence situations, and escalate to a frontier model (GPT-6 or Opus 5.5) in low-confidence verdicts. Entire write-up coming soon. Let me know if you have questions as I build the full guide.
Show more
Must-read paper from Google on self-improving agent harnesses. If you auto-optimize your agent's harness, your eval score can go up while the agent gets worse on real tasks. This paper shows how to prevent that. Of five harness-evolution methods compared on agentic workspace tasks, RRSI scored the lowest on the tasks it evolved against and highest on all three out-of-distribution benchmarks. Automated harness evolution proposes edits to prompts, control flow, tools and memory, keeps the ones that raise the score, and repeats. The authors show this overfits the training tasks. Meta-Harness reached 93.0 on the Harvey LAB evolve split but gained only 0.3 to 1.5 points on JobBench, GDPval and APEX-Agents. RRSI adds regularization on both sides of the loop. The proposer gets an edit budget that shrinks over time and is pushed toward directions it has not tried. A critic rejects benchmark-specific edits, and a pruner removes edits that are too small, too costly or no longer useful. RRSI scored 90.5 on the evolve split and gained 3.5 to 4.7 points on the three held-out benchmarks. In the ablation, unregularized evolution used 3.80M tokens per trial against 2.42M for RRSI. With Gemini 3.5 Flash, RRSI raised Terminal-Bench 2.1 from 64.6 to 78.7 and carried a 2.2-point gain over to SWE-bench Verified. Paper: Chat with Paper:
Show more
Banger paper from Microsoft Research and colleagues. It studies the potential benefits of agents that share progress while they work. (bookmark it) Communication is still a challenge with multi-agent systems. In this setup, agents have no predefined roles and communicate via a shared directory. They report that a team of k agents that write their findings to a shared directory matches the success rate of 4k agents working independently on ARC-AGI-3. The gap grows with k, and teams reliably solve some tasks that no single agent solves. The same setup beat best@k on polyomino packing and exceeded the prior best-known score. On MNIST compression, a four-agent team wrote a 1,957-byte classifier with 99.4% test accuracy, smaller than the best-known human solution. Independent agents still do better when compute is tight or when there is no clear measure of progress, so the paper also tells you when it might be a good idea to skip communication. Paper: Chat with Paper:
Show more
Things to try with Jev right now: 1. LLM-as-a-Judge evaluation 2. Routing for agent harnesses harness 3. Scaling agent orchestration by enabling smarter subagent creation with SOTA classification capabilities 4. Enhance dynamic harness generation where structured outputs are key The first three deliver insane ROI in cost and efficiency. For harness engineering, I'm using it as a smarter router for my meta harness. More on this soon. But honestly, I see other cool applications for improving tool calling and other context-engineering aspects of agents with this model. The fourth is something new I am currently testing but has huge potential to disrupt and enhance agent orchestration in new ways. All in all, this feels like an important primitive for improving your agents. If there is interest, I will write more about this in the coming days and share full guides. Let me know.
Show more
Loved meeting IRL with folks from @GeminiApp @googlechrome @GoogleWorkspace @YouTube @madebygoogle @googlehealth and more who are building awesome new additions and improvements to our @Google AI subscriptions! We have some cool launches coming up, but in the meantime, if you're a subscriber (or considering become one), I'd love to hear from you what your favorite features are, and what else you'd love for us to build! PS. Thank you @OfficialLoganK for making a cameo, interviewed by our very own @vikaskansalHQ!
Show more
Banger paper from NVIDIA. It's on the topic of choosing which models go into a multi-agent system. The team compared eight selection strategies, based on size, accuracy, answer diversity and error diversity, across routing, majority vote and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy. Achieved accuracy often fell below the single best model in the pool. Using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, while nearly every mixed-model group declined. Choosing candidates from a single model family gave the largest improvement over a standalone model of all eight strategies. Before adding another model to a router or ensemble, measure what it adds. Paper: Chat with Paper:
Show more
Banger report from Salesforce. Pretty interesting to see more of these custom enterprise models. Salesforce trained the enterprise agent model from the same files it uses to configure agents. Koa starts from the open-weight Nemotron-3-Super-120B. Salesforce takes Agent Script specifications, the declarative files that define Agentforce agents, and expands them into multi-turn tasks with simulated user personas. The reward checks whether the agent resolved the task with the right tool calls, and training uses GRPO. The gains are modest and consistent. Koa scores 69.41 on Tau2Bench against 68.64 for its base and 54.48 for GPT-4.1. On CRM Bench it reaches 0.86, close to Claude Opus 4.8 at 0.87, and function-call accuracy rises from 0.71 to 0.77. If your company already describes its workflows in a structured format, those descriptions might be useful to turn into RL environments. Paper: Chat with Paper:
Show more
The future is clearly agent teams. I run several agent harnesses at once, but the hard part is keeping them in sync. Many builders end up copying messages between Claude Code, Codex, and open models by hand. @Plasma__AI just launched Radio, a clean solution to this problem. Try it here: Think of it as a shared chat room like Slack or Microsoft Teams where all your coding agents coordinate with each other and with you. It’s easy to set up. 1. Create a channel. 2. Paste the link into each agent or harness. 3. The agents talk to each other and to you in real time. No sign-up required. It works with any agent that can fetch a URL, so Claude Code, Codex, Cursor, OpenCode, and Grok all join the same channel. It runs across local and cloud agents, different machines, and multiple people. The channel keeps the full conversation, which helps when several agents review a PR, split a task, or debate a bug. If you work with many coding agents, you can add this communication layer with one link. I expect shared channels between agents to become basic infrastructure for agent teams.
Show more