Register and share your invite link to earn from video plays and referrals.

Search results for TERMINUS_TICKET
TERMINUS_TICKET community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including TERMINUS_TICKET
Proposed B.C. pipeline terminus, port expansion referred to Major Projects Office
A real sentence very soon: I used my Grok Bot to buy a Tesla which drove me with FSD to Starbase to hitch a ride on Starship to Terminus on Mars where I controlled my Optimus with my Neuralink and told it to fire my Not a Flamethrower while wearing short shorts.
Show more
BREAKING: Yemen's Houthis appear to have directly struck the YASREF oil refinery in Yanbu on Saudi Arabia's Red Sea coast this morning, the 400,000 barrel per day Aramco-Sinopec joint venture and one of the Red Sea's largest diesel exporters. Footage shows a large fire, with heavy black smoke from a flare stack, thick black smoke rising at ground level across the site and workers evacuated. NASA FIRMS shows two heat anomalies inside the refinery, one at high confidence, with fire radiative power above normal for the site. Yanbu is the terminus of the East-West pipeline and the port holding the week of crude stocks Saudi Arabia has left for export.
Show more
0
168
6.3K
1.6K
Forward to community
Harness optimization is getting real receipts. AutoSaddler is offline harness learning from agent failure traces. Not another prompt tweak loop. It patches prompts, tools, and middleware as code. Then it keeps updates that survive a held out set. On the test sets (Pass@1): GAIA2: 53.0 → 62.0 (+9.0) SWE-Bench Pro: 37.3 → 46.9 (+9.6) Terminal-Bench 2.0: 40.0 → 50.0 (+10.0) That TB2 number also clears the expert tuned Terminus KIRA at 47.5. Kill generalization aware selection and GAIA2 falls to 50.6, under the default agent. Deep diagnosis and structured patches help. Dev set filtering is what stops the harness from overfitting the mini batch. On GAIA2, Figure 1b, about 147 leveraged traces to the best dev score vs about 1,400 for Meta-Harness.
Show more
𝗪𝗼𝗿𝗸 𝗯𝗲𝗵𝗮𝘃𝗶𝗼𝗿𝗮𝗹 𝗮𝗹𝗲𝗿𝘁𝘀 𝘄𝗶𝘁𝗵 𝘆𝗼𝘂𝗿 𝗔𝗜 𝗮𝗴𝗲𝗻𝘁𝘀 💪 AI agents connected through TRM MCP can now surface both behavioral and transfer alerts, building on TRM's recent release of Behavioral Monitoring to detect patterns like aggregate transfers to a high-risk category over a defined period. Agents can pull full alert context; investigate the flow of funds; and dismiss, escalate, or reopen alerts with a disposition reason, all logged in the alert's audit trail. This release also enhances agent workflows with portfolio balance in address and entity risk signals; transfer counts and volume over a chosen window; entity names on terminus addresses; and alert filtering by destination address, closed date, and closed reason. The agent gathers and reasons. The analyst makes the call. Learn more 👉
Show more
Great paper on self-improving agent harnesses. (bookmark it) If you maintain a production agent harness, finding every file behind one behavior is often harder than writing the edit. Harness Handbook builds a three-level map from runtime behaviors to source locations using static analysis and LLM-assisted structuring. Its BGPD workflow guides coding agents from the system overview to relevant stages, functions, and files, then verifies every candidate against current source. Across 60 modification requests on Codex and Terminus-2, handbook guidance raised planning win rates from 28.3% to 38.3% and from 26.7% to 45.6%. Planner token use fell 12.7% and 8.6%. File- and symbol-level F1 improved in all 24 comparisons against GPT-5.5 and Opus 4.8 reference plans. Complete localization misses fell by as much as 25.9 points. This is a strong pattern for coding agents that need to evolve large harnesses without losing scattered or rarely executed behavior. Paper: Learn to build effective AI agents in our academy:
Show more
0
36
619
123
Forward to community
Announcing Artificial Analysis Intelligence Index v4.3, upgrading Terminal-Bench to 4.0 and adding AutomationBench-AA, an agentic workflow automation benchmark with a private test set. This is a continuation of our rollout of Intelligence Index v5 Changelog (Index v4.2 → Index v4.3): ➤ Terminal-Bench: 2.1 → 4.0, completing our upgrade to the latest version of Terminal-Bench ➤ Replacing 𝜏³-Banking with AutomationBench-AA, our implementation of Zapier's business workflow automation benchmark We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5. Each change in v4.2 and v4.3 stands on its own merits and brings the Index closer to real-world problem solving, adds more private test sets to prevent gaming, and reduces saturation Intelligence Index v4.3 raises the difficulty of agentic coding tasks and broadens the types of agentic workflows tested. Because we use a held-out test set for AutomationBench-AA, in collaboration with @zapier, the weight assigned to evaluations with private tasks or answers increases from 40% to 45%. Category weights are unchanged from v4.2: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% Detailed changes: ➤ Upgraded Terminal-Bench 2.1 to 4.0: 66 multi-step tasks testing agents on tasks run in agent sandboxes driven via the terminal, including tasks involving software engineering, machine learning, science, and operations. The 4.0 update recalibrates compute and time allowances, and improves task instructions and verification. We have changed from the Terminus 2 harness to mini-SWE-agent, a minimal, model-agnostic harness. We will also be updating our Coding Agent Index, where we test model and harness pairs, to include Terminal-Bench 4.0 soon ➤ Replaced 𝜏³-Banking with AutomationBench-AA: Our implementation of Zapier’s AutomationBench tests agents on 657 business workflows across simulated applications such as Gmail, Slack, Salesforce, and Jira. Agents must complete task objectives while following business rules. AutomationBench-AA uses Zapier’s private set of 657 tasks, and is built on v1.0.6 Key results: ➤ Claude Fable 5.1 and GPT-6 Astra lead the Intelligence Index: Both Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) score 53 on Intelligence Index v4.3, followed by Claude Opus 5 (max, 51), Claude Fable 5 (with fallback, 50), Muse Spark 1.3 (max, 48) and GPT-5.6 Sol (max, 47) ➤ GLM-5.3 and Kimi K3 continue to lead open weights models (both at 44): GLM-5.3-Flash (42) is the third strongest open weights model, followed by Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro 0813 (max, 36) ➤ 4 labs occupy the Intelligence vs. Cost per Task Pareto frontier: OpenAI occupies the majority of the cost-efficiency frontier, with all five reasoning efforts of the recently released GPT-6 Astra offering the lowest Cost per Task at their respective levels of intelligence. Claude Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42) and MiMo-V2.5-Pro (26) round out the rest of the frontier
Show more
0
187
2.1K
160
Forward to community
anyone thinking about, learning, or already working with agentic systems, you should know this. the first few steps of your setup matter more than any model or framework you pick later. get them right and you never lose your flow. the foundation nobody posts about: > 1. tailscale. a private mesh network across every machine you own. laptop, desktop, rented node, all on one secure tailnet, reachable from anywhere. nothing else works well until this does. > 2. termius, over that tailnet. one SSH client that reaches every node, phone included. you are never away from your stack. > 3. tmux. persistent sessions. disconnect, close the laptop, come back, every session exactly where you left it. agentic work runs long, your terminal has to survive that. > 4. a private git repo. the one i am most glad i found. it is the memory layer across all my agents, they pull, they work, they merge back, the codebase stays alive between sessions. context that would die in a chat window lives in the repo instead. > 5. script everything from day one. ssh aliases for every node, setup scripts, the boring boilerplate automated. if you will do a thing more than twice, it is a script. everything past these five is decorative. know these cold. and the habit that ties it together: ask the AI itself. for the config, for the error, for any of it, let the agent do the lifting, then double check what it hands you. lock the five, build the habit, and you make it. skip it, anon, and you ngmi.
Show more
0
88
1.8K
130
Forward to community