Register and share your invite link to earn from video plays and referrals.

Composio
@composio
Your agent is smart. Its tools should be too. Check
49 Following    24.5K Followers
We tested 6 AI models on 30 challenging agent tasks: GPT-6 Astra, Opus 5.5, GPT-6 Sol, Pareto 26.9, DeepSeek V4 Pro, and GLM 5.3 Flash. Sol matched Opus’s score, finished faster, and cost about a quarter as much per successful task. Here’s how all 6 models compared 🧵🧵🧵
Show more
Here's how all 6 harnesses compared on task success rate: Codex, Hermes Agent, and Command Code tied for the lead at 21/29. Claude Code was one task behind, while OpenCode and Pi Agent finished at 19/29. The spread was about 7 percentage points. Only 3 tasks produced different results across harnesses: 18 passed across the board, and 8 tasks also failed across all harnesses.
Show more
We just added support for Pinterest and Pinterest Ads to Composio. This means you can now connect Pinterest to your agents and do things like: → Find fast-growing shopping categories → Analyze your top-performing Pins → Search your Pinterest content → Manage boards and Pins → Pull Pinterest Ads performance → Inspect campaigns, ad groups, ads, and targeting → Pause or resume ad campaigns Plug Pinterest into Claude, Hermes, ChatGPT, Grok Bot, or your favorite agent.
Show more
Celebrating a new product release with one of our special customers🥳🤫👀 (with some non-alcoholic wine because tech-bros don't drink)
Inside @Stripe's MPP IRL: machine payments, stablecoins, and the future of agent commerce. Agents are starting to pay for things on their own: API calls, services, each other. Nobody's agreed on how that should work. This panel was three takes on it. @BackseatVC (Stripe) leads product for Machine Payments — the open standard for how agents pay for things. @dwr (@tempo) is building stablecoin payment rails, and thinks that's what agent commerce runs on, not cards. @blauyourmind (Royal) magician turned @a16z Crypto partner turned CTO is exploring programmable money for creators, and why the demand side is barely here yet. We got into: • Why HTTP 402, a status code from the early web, is suddenly the backbone of agent payments • Why Dan thinks a credit card is a private key and why stablecoins are safer for agents • Where stablecoins actually win first • Whether you should build for agent payments now or wait • Why the demand side is "virtually nonexistent" and what that means if you're building 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Event Kickoff JEN LEE — Product Lead, Stripe (00:44) What the Machine Payments Protocol Is (01:28) How MPP Works — the HTTP 402 Payment Challenge (03:25) Adoption So Far: ~30,000 Transactions (04:04) Building Trust — Shared Tokens & the Link Agent Wallet (05:56) What the Creator Economy Looks Like When the Audience Is Agents (07:15) What She Was Certain About at 22 That's Now Wrong DAN ROMERO — GTM, Tempo (08:34) Meet Dan (08:49) Why Stablecoins, and the Genius Act Tailwind (09:48) "A Credit Card Is Basically a Private Key" (11:19) Should Every Company Be Thinking About MPP? (13:03) What to Be Wary Of (15:15) Where Stablecoins Win — Payouts, Remittances, DoorDash MICHAEL BLAU — CTO, Royal (16:36) Meet Michael (17:13) From Magician to a16z Crypto to CTO (17:57) Team Over Idea — His a16z Takeaway (18:26) "The Demand Side Is Virtually Nonexistent" (19:03) Closing
Show more
We tested DeepSeek V4 Pro 0813 across 5 different agent harnesses on 30 challenging agentic tasks. We compared Claude Code, DeepSeek Harness 0.1, Hermes, Pi Agent, and OpenCode. Pi Agent solved the most tasks, while DeepSeek’s own harness won on cost. 🧵🧵
Show more
We just added 15 new apps to Composio today: - @toggl — Time tracking for teams and projects - IPLocate — IP geolocation and threat intelligence - @BulkEmailCheck — Bulk email verification - Sideshow — Visual workspace for coding agents - @Exceptionless — Error monitoring and reporting - Company URL Finder — Find company domains from names - @debouncer — Email verification and list cleaning - — Email validation and deliverability checks - @vast_ai — Marketplace for cloud GPU compute Cronfree Time Scheduler — Schedule recurring automated tasks - @_magicslides — Generate presentations with AI 44API — VAT and tax ID validation - @handwrite_io — Send automated handwritten cards - Invalid Bounce — Detect invalid and risky emails - @ConsensusNLP — AI search for academic research We’re making your agents more powerful with more tools, day by day.
Show more
We tested DeepSeek V4 Flash against Kimi K3 and GLM 5.2 on 30 challenging agentic tasks. DeepSeek was 2.5x faster than GLM and 1.4x faster than Kimi at a similar success rate, despite using the most tokens per task. 🧵🧵🧵
Show more
Hermes and Pi Agent led on the average cost per task, while Claude Code cost about 3.7x as much as Pi: - $0.39 Hermes Agent - $0.40 Pi Agent - $0.47 Codex - $0.51 OpenCode - $0.54 Kimi Code - $1.47 Claude Code The median cost tells the same story: $0.29 in Pi Agent and Hermes, $0.35 in OpenCode, $0.38 in Kimi Code, $0.39 in Codex and $0.72 in Claude Code, so the cost gap holds for a typical task and is not driven by a few expensive runs. We calculated these costs using Kimi K3’s list prices: $3/1M input tokens, $0.30/1M cached input tokens, and $15/1M output tokens.
Show more
Lesson: If you want to reduce agent costs, examine the harness before switching models. In our data, the harness changed the cost by 9× while the model’s capabilities remained about the same.
As we mentioned, success rates stayed close, but speed and token efficiency split between harnesses: - Success rate: 22/28 for Kimi Code, 21/28 for Hermes, 20/28 for Claude Code - Median time per task: 179s in Hermes, 297s in Kimi Code, 348s in Claude Code So the fastest harness (Hermes) and the most token-efficient one (Kimi Code) were different.
Show more
The median task consumed nearly 6x more tokens in Claude Code than in Kimi Code: - 61k in Kimi Code - 67k in Hermes - 340k in Claude Code At K3's $3/M input rate (input tokens make up roughly 95% of agentic workloads), the average cost per task was: - $0.22 in Kimi Code - $0.28 in Hermes - $2.00 in Claude Code
Show more
We ran Kimi K3 through 3 agent harnesses (Claude Code, Hermes, Kimi Code) on 28 identical tasks. All 3 harnesses completed the tasks at similar success rates, but the interesting story is token efficiency: the same task cost up to 30x more tokens depending on the harness. 🧵🧵
Show more
0
210
2.7K
157
Forward to community