Register and share your invite link to earn from video plays and referrals.

Einsia
@EinsiaAI
We're teaching AI the work the world's experts actually do.
6 Following    699 Followers
An AI agent spends hours on a task. Why should all that work disappear when someone else takes over? Einsia AI’s answer is AgentGit—an open-source platform for collaborating on agent sessions, so work can be saved, handed off, and continued by the next person. Explore how others solve problems, and share your agent experience with the world. Try AgentGit 👇 #OpenSource# #AIAgents# #DeveloperTools# #DevTools#
Show more
0
135
612
96
Forward to community
The ultimate test for coding agents isn't local editing— it's whole-repo evolution, and right now, the survival rate is 5.4%. Today we’re releasing SWE Refactor Bench, a benchmark for long-horizon, whole-repository software stack migration. Coding agents are getting very good at fixing bugs. But can they refactor an entire system, C → Rust, Maven → Gradle, POSIX → WebAssembly? We built 20 real migrations across projects, including SQLite, zlib, libsodium, and GraphHopper. 520 runs. Only 28 survived all 3 stages. 13/20 tasks were solved by nobody. System-scale migration is still wide open. Full breakdown 👇 GitHub: [ Paper Link: [ Einsia Website: [
Show more
0
98
2.3K
93
Forward to community
1/ Recursive self-improvement (RSI) depends on agents improving how AI systems are trained —not just tuning hyperparameters, but improving the training algorithm itself. We tested this directly with AI4AI-Bench: 10 real research repositories spanning 10 distinct algorithm families. Full breakdown 👇 GitHub: [ Paper Link: [ Einsia Website:[ 📊 The results: The average score is just 0.166. Even the best-performing model, Opus 5, reaches only 0.288. The median exploration cost per task rises from $1.69 to $34.60. #AI4AI# #RecursiveSelfImprovement# #AIResearch# #AI4AI_Bench#
Show more