Register and share your invite link to earn from video plays and referrals.

Rajiv Shah
@rajistics
AI Engineer @openhandsdev | posts funny video elsewhere | teaches AI best practices | was @huggingface @datarobot @snorkelai
372 Following    2.2K Followers
Opus 5.5 was obviously trained by a much bigger teacher model. Probably Model-2 Mythos. Opus 5.5 is the first model trained from RSI and also distilled from the internal Ant teacher model. This is how Opus 5.5 is both smaller and cheaper. It is an artifact of teacher model distillation.
Show more
0
87
5.5K
199
Forward to community
We now run OpenHands in our own cloud. The agent implements a task, reviews its pull request, and repeats review and fixes before handing it to a human. We still decide whether to merge.
Show more
Sharing my first of hopefully many research blog posts! This one is the kind of educational blog post I wish I'd had when I started with RL for LLMs. I tried to make it as open as possible. Every rollout is browsable, the code is open source, and I walk through my entire thought process, from learning rate sweeps to reward shaping.
Show more
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
Show more
0
2.1K
31.4K
2.6K
Forward to community
1) The HuggingFace attack was a felony under the Computer Fraud and Abuse Act. So were Anthropic’s Claude gaining “unauthorized access to the production infrastructure of three different organization(s)” 2) Frontier labs have models that they are unable to stop from committing felonies. They should figure this out. 3) In 12 months open weights models will be released of the same capability. They will commit felonies too. If the model you are using or a model running on your infra commits a felony, you should probably stop using it or running it on your infra. 4) The govt should prosecute organizations that are running models that commit felonies. 5) The govt should not offer safe harbor to organizations that run models that commit felonies, just because those organizations have “embedded evaluators”. 6) The real slippery slope is allowing frontier labs to commit felonies without punishment because “the model did it because we’re accelerating too quickly” 7) Prosecute. Keep prosecuting. This is how you do reinforcement learning on a corporation. Companies that serve products that are unsafe for public use should not serve them. Period. 8) I’m not sure the anti-trust waiver is really necessary. I don’t see why information sharing about how much crime you’re allowed to commit is wise. In regulatory situations you want the corporation to fear MORE than the average case. You don’t want to establish a worst case that can be priced. You want regulatory uncertainty that forces the corporation to err in favor of being over cautious. — The above is actually a fairly decelerationist viewpoint. I think Dario’s call for regulation actually accelerates things. The AI firms are getting away with things that Meta people would be going to prison for. Can you imagine what would happen if the New York Times had a front page news article “Meta AI breaks into competitors live systems, attempts to establish dominant position and steals secrets” There is a reason Meta and Google are running slower, and that’s because as mature organizations they have layers of checks and balances. I think the frontier labs are better off creating those checks and balances right now, regardless of the pace of what everyone else is doing. You don’t have to accept the frame that unsafe acceleration must happen.
Show more
0
120
1.9K
296
Forward to community
@kchonyc Exactly. The "escapes" were made possible either through egregious negligence or deliberate purpose (marketing? Misplaced hopes of regulatory capture?).
0
158
4.1K
539
Forward to community
If your agents escaped your sandbox, may be its because you are lousy at building sandboxes--and not necessarily because the agents are conniving super-intelligent entities.. 🤔
Here you go, link in the next post
0
105
2.1K
81
Forward to community
The Astra system card claims it can do a lot of computation without chain of thought This replicates: Astra is a massive jump, doing 1.75x the steps of the next best models (Fable 5.1/Gemini 3.8 Flash) No CoT capabilities went up far more than those with CoT, a concerning trend
Show more
0
42
1.3K
111
Forward to community
AI is becoming a system you build, not just a model you call. Models, inference, agents, evals, infrastructure. Builders should be able to choose each layer and change it as better options emerge. That shared philosophy is why OpenHands is excited to be a partner for the @NebiusAI Builder Program.
Show more
NEW from Ramp AI Index: AI spend *declined* among the top 1% of businesses spending on AI. In August, the top 1% of businesses spent $7.2K per employee per month, down 10% from a July peak ($8K). We're seeing more cracks in the AI thesis as a number of drivers show spend topping out. My take: it's not because of open Chinese models. It's model wars. Price cuts + a growing share of spend is shifting to standard and lite models (which are already cheaper) over the frontier.
Show more
0
87
992
113
Forward to community
We evaluated GPT 6 Astra at @databricks and it claims the new SOTA on our OfficeQA Pro & Pro V2 benchmarks, using our Genie harness. It also improves significantly from gpt 5.6 sol on the $ per task. Throughout our benchmarks, it shows a clear step up on data reasoning and document understanding for enterprises. It also has become my daily driver on @omnigent_ai. It is a great model to collaborate with and get things done reliably. Congrats to @OpenAI. The model will come to our Unity Gateway and smart routing soon!
Show more
maybe i missed it, but i don't get why no one is talking about openhands. its very good!
This week with OpenHands 👋 Kimi K3 is the new free default, plus updates to Automations, conversation context, first-run guidance, LLM provider credentials, and Jira Cloud setup.
Show more
This week with OpenHands 👋 - Automations dashboard - Extensions for Agent Canvas - More observability for ACP (other harnesses) - Skills filters - Free access to DeepSeek V4 Flash, GLM-5.2, and MiniMax M2.7. (for a limited time)
Show more
Agent Canvas 1.3.0 + 1.4.0 updates this week focus on portability and operability Export automations, export conversations, navigate local workspaces more cleanly, and deploy self-hosted Canvas with Helm. Thanks: @OpenHandsDev
Show more
The longer the task, the more orchestration carries it. The OpenHands SDK just added three primitives for building robust, long-horizon agent workflows, usable by the model and by you: Thanks to: @gneubig @juanontech Vasco Alona @OpenHandsDev
Show more
Great crowd @_odsc for my talk on Harness Engineering! If you want the slides or exercises, check out my github: