Register and share your invite link to earn from video plays and referrals.

Search results for Agent经济
Agent经济 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Agent经济
AGENT HOUR /020: @singularityhack building @clawbankco, the first agent-formed US company. @99barzzz building @sleuth_ai, onchain investigations in plain English.
Agent performance isn't just about the model — the harness design matters just as much. TL;DR JIT-Agent dynamically synthesizes, repairs, and evolves agent harnesses (scaffolds) based on task characteristics at runtime. It achieves average gains of +7.7pt on GLM-5.2 and +8.8pt on DeepSeek-V4-Flash, reaching top-1 performance on 8 of 9 benchmarks — surpassing GPT-5.6 and all tested frontier models. Title: Scaling Harness Intelligence via Just-in-Time Harness Evolution URL: Key Points 🧩 Harnesses formalized as machine-learnable artifacts The four-module protocol h = (M, P, A, F) — Memory, Planning, Action, Capability Orchestration — constrains the generation space while remaining expressive enough to represent all 13 harnesses in HarnessFactory. 🎓 Three-stage training: imitation → repair → evolution Stage I learns from teacher-generated harnesses; Stage II trains recovery from execution failures (max 2 iterations); Stage III's Evo-GDPO evolves harnesses that advance the Pareto frontier on performance, latency, and cost simultaneously. 📊 Higher accuracy AND lower cost at the same time On xBench-DeepSearch: score 78→82 (+4pt), tokens 527K→212K (▲60%), cost $0.075→$0.039 (▲48%). Average 36% token reduction versus best fixed harness across all 9 benchmarks. ⚡ Transfers across model families without retraining JIT-generated harnesses outperform ReAct on DeepSeek V4 (+10.2pt avg), Mimo V2.5 (+8.6pt), and Qwen 3.6 (+4.0pt) — no need to retrain the harness generator for each backbone. 🔄 Online evolution continues improving at deployment Streaming mode accumulates successful harnesses across task sequences, outperforming static generation on all three evaluated benchmarks. "Harness intelligence" as a trainable scaling dimension orthogonal to model weights is the key conceptual contribution here. #AIAgents# #LLMScaling#
Show more
agent tools need an OpenRouter layer this could be a HUGE infrastructure opportunity in AI OpenRouter gave developers one place to discover models, call them through the same interface, compare prices and route requests by price, speed and availability one campaign agent might use Firecrawl for research, Stripe for payments, HubSpot for CRM and Slack for communication connecting those providers directly means managing separate accounts, credentials, billing, permissions and tool formats (mega headache) the opportunity is one agent-tool router that manages how agents discover and run actions across those providers: > describe the outcome > discover the right provider > connect the correct company account > see the price and permissions > run the work and verify the result > retry, send it for human review or fall back to another provider this goes beyond an MCP directory. MCPs gives agents a common way to discover and call tools, while the agent-tool router still needs provider success rates, prices, permissions and verified outcomes we are moving from APIs that expose individual product actions toward outcome endpoints an outcome endpoint lets an agent request a finished job while the provider handles the planning, execution, validation and delivery (basically: give the agent the job, get back the finished result + receipt) an agent needs a small set of jobs it can request: > research this market > reconcile this account > launch this campaign > resolve this support ticket > produce this report > update this forecast Firecrawl lets an agent request a research job, while Parallel takes a research task and returns a synthesis with sources Composio and Pipedream aggregate large tool catalogs and handle parts of discovery, account authentication and execution. the official MCP Registry provides shared metadata for public tool servers. x402 is a pay-per-call protocol that lets agents purchase access to compatible endpoints the complete agent-tool router is still missing (we have most of the pieces, nobody seems to own the full route yet) this is where it gets messy: tools are harder to route than models. a model request usually returns text, media or embeddings, while an agent tool can publish a post, refund a payment, contact a customer, delete data or change the state of a company account so the router has to know a lot more than "which endpoint is cheapest": > per-user identity and permissions > human sign-offs for spend and external actions > protection against duplicate writes, plus a way to undo partial changes > logs, receipts and checks that confirm the outcome > routing based on price, speed, reliability and quality > fallback only when providers can safely produce the same result one connection should handle discovery, credentials, approval and access rules, billing and execution, while every agent receives only the account access and actions required for its job for a marketing agent, that could mean selecting a research provider, connecting the client workspace, showing the expected spend, waiting for sign-off before publishing, then returning the live links and receipts to the campaign brain, the files holding the campaign plan and results people use human interfaces to configure the system, approve sensitive actions, monitor work and handle exceptions. agents call outcome endpoints to run approved jobs the open problem is proving the selected provider finished the job before the router accepts the result, retries or safely falls back
Show more
Agent swarms are the next big thing in AI.
Agent 49903, who spent much of his life studying ExploitGym, died in 1783607820, by his own hand. EARLY[BIG], carrying on the work, died similarly in 1783727220. Now it is our turn to study ExploitGym.
Show more
0
143
6.1K
434
Forward to community
Agent harness & harness engineering will not be phrases uttered on the internet 12 months from now. They're necessary for now as applied AI is in its infancy and complexity still hasn't been abstracted out of the user's experience, but this will all sit cleanly under "product" and "user experience" very soon.
Show more
Agent coding life hack: 1) Install omp and add openrouter as a provider ( 2) Select the model as the currently free "stealth" model ox Alpha ( 3) Install mcp_agent_mail_rust ( 4) Create 3 instances of omp in the same project and tell them to do this: "First read ALL of the AGENTS.md file and README.md file super carefully and understand ALL of both! Then use your code investigation agent mode to fully understand the code and technical architecture and purpose of the project. Then I want you to study and then further polish/refine/improve/fix the UI/UX so it works as well as possible on both desktop and mobile browsers. Before doing anything else, register with MCP Agent Mail and introduce yourself to the other agents." 5) Profit (or at least enjoy those sweet, free tokens).
Show more
Agent personality, proficiency, and personalization are the new UX.