Register and share your invite link to earn from video plays and referrals.

Sentient
@SentientAGI
To ensure that Artificial General Intelligence is open-source and not controlled by any single entity. @SentientEco @OpenAGISummit
69 Following    535.2K Followers
Catch Sentient researcher @TankDarshan7 at @kdd_news 2026 to learn why LLM agents struggle with multi-tool, long-form analysis. Whether you're a researcher, builder, or working on LLM agents, we hope to see you in Jeju, South Korea 🇰🇷
Show more
“CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis” has been accepted as a poster at KDD 2026 (@kdd_news) — one of the leading conferences for AI, machine learning, and data mining. At KDD? Meet with Sentient researcher Darshan Tank (@TankDarshan7) to learn why LLM agents struggle with multi-tool, long-form analysis in crypto research.
Show more
Closed labs can watermark outputs because they control the endpoint. Open weights have no such control. That’s why our research embeds fingerprints directly into the model. It is designed to survive fine-tuning, merging, and distillation. The mark belongs in the weights, not the output.
Show more
Anthropic says new Claude models will embed invisible watermarks in all generated text, everywhere Claude is offered. The watermark is part of the text, it isn't metadata: "it will travel with the text when it's copied and pasted elsewhere, and may persist through some editing." This starts with models launched on or after August 2, 2026, under an EU AI Act code Anthropic signed. Anthropic is still working on adding it to current models. The rollout is worldwide.
Show more
Every word on this page is a real AI term, except for one we made up. No AI has ever been trained on it. No paper has ever used it. Can you spot which one is the fake?
This is why checking your agent's math isn't enough ↓
Most AI agents don't just get the math wrong. They get the number from the wrong place: outdated table, mismatched row, wrong year. That's why @abraxasnz13 built an agent that won't calculate until it knows where the number came from ↓
Show more
Proof that a better harness won't crack a harder task ↓
You can't benchmark your way around a hard problem. Across 13,000+ agent runs in the OfficeQA Public Challenge, Sentient researchers @iamnamanvats and Deep Halder found that different harnesses agreed 88-93% of the time on which tasks succeeded and which failed. TLDR: Difficulty lies in the task, not the harness.
Show more
Turns out our runner-up's best piece of prompt engineering was a backspace ↓
One of the best decisions @Fr0oZi made during the OfficeQA Public Challenge was to delete a single line: "You are a financial analyst." Here's how he placed 2nd in the Arena using @goose_oss and @MiniMax_AI M2.7 ↓
Show more
From a hunch to hard data, all in the open ↓
Some come to Arena to compete, but Cat Varnell (@hypecatv2) came to run an experiment: treat AI as a collaborator, not a tool, and it performs better. She left with data to back her theory, and a community that shares her vision for open-source AI ↓
Show more
Proof that open-source AI summer has arrived ↓
The first open 3T model releases its weights, NVIDIA CEO debuts on X in support of an open AI ecosystem, and 36+ companies join the Open Secure AI Alliance. Open-source AI summer has arrived ↓
Show more
ICYMI: Run self improving coding agents without Claude API and on @FireworksAI_HQ Read the full EvoSkill v1.3.0 breakdown by @AlphaSignalAI:
Here's what a 1st place build in the Arena looks like ↓
The OfficeQA Public Challenge wasn't won with a smarter model. It was won with a system that made it harder for the model to fail. @_MAYANKMAHAUR_ shows you how he did it ↓
Show more
What makes an agent run succeed vs fail? We pulled 13,483 runs across 156 submissions and 4 harnesses from the OfficeQA Public Challenge to find out ↓
Closed labs are building a future you'll never see inside of. We're putting $42M behind the one you can ↓
From NeurIPS to ICLM and beyond. Explore every paper from Sentient researchers, all in one place:
This is how open-source AI turns competitors into teammates ↓
For Shubham (@musickeeda), Arena meant sharing a room with other teams and chasing better results together, even as competitors. That's what open-source AI hackathons make possible ↓
Show more
Our EvoSkill research paper has been accepted to Agent Skills '26 at @CAISconf, a workshop dedicated to the design, evaluation, and optimization of skills for LLM agents. Led by @salahalzubi401 and members of the @virginia_tech research team, EvoSkill is an open-source toolkit that takes a benchmark and a coding agent, and evolves it into a state-of-the-art specialist in minutes.
Show more
Want to start building your own self-evolving agents? @0xhermes_ shows you how to set up EvoSkill from scratch using the same workflow that helped him place 2nd in the Arena ↓
Show more
EvoSkill has now been cited by 56+ papers. From frontier AI labs at Microsoft Research (@MSFTResearch) and Tongyi Lab (@Ali_TongyiLab), to top universities such as National University of Singapore (@NUSingapore) and Columbia University (@Columbia), researchers around the world are building on our work that helped pioneer the field of self-evolving agents. Proof that open research travels fast.
Show more
Fireworks AI is now live on EvoSkill v1.3.0! You can now use @FireworksAI_HQ directly with EvoSkill to run fast inference on open models as both the evolution harness backend and the LLM scorer. Alongside Claude API and OpenRouter, Fireworks AI is now a first-class provider in EvoSkill.
Show more
Which open source AI use case should @sentient_found fund?