Register and share your invite link to earn from video plays and referrals.

Sentient Ecosystem
@SentientEco
Empowering @SentientAGI builders, researchers, and the ecosystem in advancing open-source AGI 🤝
74 Following    28K Followers
Errors aren't a red flag for agents. Across 13K+ OfficeQA runs, both passing and failing agents hit errors at nearly identical rates. TLDR: An error isn't a sign the run is doomed, so counting errors is a bad way to predict failure.
Show more
Compute buys capability. But it doesn't tell anyone where the model came from.
Models in 2026: All accusations, no fingerprints. Mark the model, end the argument.
When an agent takes more actions but makes no real progress, this is a useful early-warning sign that it’s heading towards failure. Across 13K+ OfficeQA runs, agents that ultimately answered incorrectly took up to 50% more steps, consumed ~40% more compute, and incurred ~40% higher cost per episode than agents that reached the correct answer. TL;DR: Failed trajectories don't just take longer. They waste significantly more compute, tokens, and money.
Show more
There’s no such thing as a problem.❌ Problems are opportunities for co-creation & change.✨ The people, wise in the ways of #CivicAI🌱#, are truly the superintelligence! Watch my Open Commons appearance with @sentient_found’s Sachi Kamiya.🙏 ▶️ #LLAP🖖#
Show more
“CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis” has been accepted as a poster at KDD 2026 (@kdd_news) — one of the leading conferences for AI, machine learning, and data mining. At KDD? Meet with Sentient researcher Darshan Tank (@TankDarshan7) to learn why LLM agents struggle with multi-tool, long-form analysis in crypto research.
Show more
Most AI agents don't just get the math wrong. They get the number from the wrong place: outdated table, mismatched row, wrong year. That's why @abraxasnz13 built an agent that won't calculate until it knows where the number came from ↓
Show more
You can't benchmark your way around a hard problem. Across 13,000+ agent runs in the OfficeQA Public Challenge, Sentient researchers @iamnamanvats and Deep Halder found that different harnesses agreed 88-93% of the time on which tasks succeeded and which failed. TLDR: Difficulty lies in the task, not the harness.
Show more
One of the best decisions @Fr0oZi made during the OfficeQA Public Challenge was to delete a single line: "You are a financial analyst." Here's how he placed 2nd in the Arena using @goose_oss and @MiniMax_AI M2.7 ↓
Show more
On Monday we showed Qwen3.7 Plus trading Bitcoin Perps through 13 of the wildest stretches in BTC history. It got liquidated in 3 of them. Then it read the records of its own deaths and rewrote its own trading rules. No human edited anything. Now it survives all 13. This is @SentientAGI EvoSkill: a skill file that improves itself from its own results 🧵
Show more
Our open-source meta-agent framework research was grounded in the premise that there should be an intuitive structure out there in the open, general and flexible enough for a wide variety of use cases, and transparent enough to let creators and users debug and understand the flow very well. ROMA is built on hierarchical recursive decomposition, which lets one prescribed topology be reused across many use cases and removes the need for anyone to rebuild the scaffolding around the agents from nothing. "At a very high level, information in ROMA flows top down through decomposition, bottom up through aggregation, and left to right through conditional dependencies between sibling subtasks. You start off with this root node, which is the original task, and it goes through top-down decomposition. It passes through the atomizer, whose job is to decide whether the task is simple enough for a normal LLM to execute, or complex enough that it needs to be broken down and planned out. If it's atomic, it goes straight to an executor. If it's not, the planner generates different subtasks, and these subtasks should ideally work towards achieving the final goal. Some of them are conditioned on the ones before them, so they can't really run until those finish, and the ones that aren't can all run in parallel. Then once the subtasks are completed, you have bottom-up aggregation, where the parent node looks at all the results of these nodes and tries to resolve its original task using those results. And this is the same logic for every node, which is what makes the whole thing recursive." ROMA dictates a standard topology for many use cases, and we're claiming this topology is interpretable, composable, and very easy to iterate on top of. We're taking the structural guesswork out of building these systems, which opens doors for enthusiasts, explorers, and builders, because they can vibe-prompt their way into whatever specific application they want without worrying about structure or setting things up. All they have to worry about is how the aggregator aggregates information or how the planner plans, and these are essentially just agents with particular prompts. Given the reduced barriers to entry, they can contribute, explore different use cases, and experiment with many different tools and mix them together. Anyone can plug in a code interpreter, a bug finder, a video generator, an image generator, a comic book generator, or whatever they find useful.
Show more
The OfficeQA Public Challenge wasn't won with a smarter model. It was won with a system that made it harder for the model to fail. @_MAYANKMAHAUR_ shows you how he did it ↓
Show more
Sentient Sparks @kurlyk27 and @Fr0oZi discuss how they use EvoSkill for their own agents ↓
The next AI system will be built by the one before it. AI is automating its own engineering lifecycle, and whoever controls that loop controls what comes next. This is why it needs to be open. @qzxcle breaks down a talk by our Product Lead @oleg_golev on how automating AI engineering can build agents that reason reliably for everyone.
Show more
@oleg_golev on automating AI agent engineering
Say hi to our friend Travis Good (@iridiumeagle) 👋 When he’s not building high scale machine intelligence as CEO of Ambient (@ambient_xyz), he's speedwalking, scrolling X, or using the word “contumacious”.
Show more
Cocktails, jazz, and 300 of SF's best AI researchers and engineers in one room building together 🍸 @sentient_found is grateful to be a part of this night with our friends from @smallest_AI and @ThineAI.
Show more
Our Chinese BD Lead @Anitahityou will be on a panel with @MiniMax_AI at @themu_xyz Shanghai to talk about EvoSkill and the AI consumer ecosystem. 📍 Alibaba Hongqiao 🗓️ Thursday, May 14th, 7-10 PM CST RSVP:
Show more