Register and share your invite link to earn from video plays and referrals.

Junyang Lin
@JustinLin610
building p7k | ❤️ 🍵 ☕️
2.1K Following    89.9K Followers
a life update: i started a new company called Pragmatik (p7k) Labs (语用科技) in shanghai, focusing on the research of next-generation agents across digital and physical worlds. thanks to Gaorong Ventures and HSG (红杉中国 & 高榕创投) for co-leading this round, and to Tencent (腾讯) and Shanghai Engine Fund (上海未来产业基金) for the support. @pragmatik_labs ·
Show more
0
219
1.7K
139
Forward to community
nice work and great to see new arch working! btw the small model is not opensourced?
Today, we are introducing Inkling. Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available. Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Show more
1/ Still looking for a minimalist, high-performance framework for agentic RL research? Meet Molt — an agentic-first, PyTorch-native reinforcement learning framework with roughly 9K lines of RL code for 700B models. ⭐
Show more
i recently discussed with some friends about the strike of model effects and costs. while frontier models are smashing a lot of things it is always expensive to use them. me personally always think about using as many tokens as possible with my subscriptions and coding plans, and i think if i turn the usages to tokens that might be a much bigger amount of money (i remember a media previously said if u use codex crazily u can use over 8000 dollars a month with 200 dollar subscription. not sure about the number). how about enterprises that really use tokens and count expense on the unit of token? some interviewees told me that they tried their design their system by tasks, like easier or to say more formatted tasks go to seed, qwen, dpsk, harder tasks like agentic tasks go to fable or sol. glm 5.2 is nice for coding but still after claude and gpt, recent grok is a bit similar. this just reminds me that if we design the system well on the user side, we can have more productivity based on sufficient usage of ai today without too much worry about expense. maybe there is still much room for this.
Show more
anything interesting recently in the opensource community?
cool
Bridgewater used their unique financial knowledge and partnered with us on @tinkerapi to fine-tune a model that helps their analysts focus on what's important. Experts improving AI that empowers experts.
Show more
wen people say sd they mean seeddance instead of stable diffusion. f hard for me to react on this
i still remember the discussion of the case of us visa application and i thought that must be mission impossible for ai and just said u guys must be crazy... now it seems that i am the dumb hahah! anyway, time for frontier models to fight again! bon courage!
Show more
Two years ago, we built OSWorld 1.0 — the benchmark that became the standard for computer-use agents. Agents now score 83.5% on it. Problem solved? Not even close. 🚀Today we introduce OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks. What's new: 🎯 108 real-world workflows, each ~1.6 hours ⏱️ for a skilled human ⚙️ ~318 tool calls/task vs. ~30 in OSWorld 1.0 🌍 Grounded in authentic artifacts & stateful user profiles ⚡ Captures real phenomena: dynamic environments, streaming interaction, cross-source reasoning, implicit-state inference & more 📊 Best results: Claude Opus 4.8 reaches the highest accuracy at 20.6%, while GPT-5.5 is far more token-efficient but plateaus near 13%. No one is close to solving real computer use. 🏠 Homepage: 📄 Paper: 💻 Code: 🤗 Dataset: 🧵 [1/8]
Show more
y codex so dumb these several days
ai making ai better recursively is an impressive idea but sometimes depressing as human researchers seem to become less and less significant in this process. i do believe that there is great potential in this direction while human supervision will become more principled and higher-level. then imagination and vision matters more than ever.
Show more
what do u think of gemma 4 12b?
Today, we’re releasing Continual Learning Bench 1.0: the first, realistic benchmark for measuring how AI systems can improve in online settings. Benchmarks today assume models are stateless. Each example is independent, and once a system finishes a task, it moves on as if nothing happened. But deployed AI systems should learn from experience. We tested 10+ frontier systems against novel, expert-validated tasks and find there’s still plenty of headroom for learning. (1/n)
Show more
0
42
1.2K
167
Forward to community
why preview
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at via Expert Mode / Instant Mode. API is updated & available today! 📄 Tech Report: 🤗 Open Weights: 1/n
Show more
i do like this passage, and here are some thoughts: 1. critical thinking is essential in the era of agents. i still remember that many years ago when i studied the lesson of critical thinking, i learned that keeping debating with yourself by listing out reasons can really deepen your thinking. today, critical thinking becomes humans debating with agents, so that they can think more deeply together and analyze problems in a more comprehensive way. 2. designing a healthy and well-structured organization and system is essential for creation and building. with systematic support and efficient tooling, humans can work exponentially more effectively together with agents. that gives people more time to take care of their physical and mental health, while also exploring new opportunities. 3. new era often favors newbies, because they have less past experience and therefore less fear of current difficulties. what oldbies should really think about is which parts of their experience are actually worth leveraging. from my perspective, we should think more carefully about which experiences are truly aligned with first principles. but anyway, ai first is super, super exciting!
Show more
we need agent evals that are really consistent with real world usages. otherwise people are optimizing foundation models for the wrong direction. the problem of targeting is even bigger than benchmaxxing.
Show more
🧵 1/ Our agent Terminator-1 scored ~100% on 8 major AI agent benchmarks, e.g., SWE-bench Verified & Pro, Terminal-Bench, beating Claude Mythos. It solved 0 tasks. Benchmarks are the field's shared language for measuring AI progress. Our new work shows that language is broken. Here’s how.
Show more
happy horse is insanely happy
🚨 Happy Horse First Output This model beats seedance 2 on artificial analysis for more information check quoted tweet
unbelievable...
Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans.
Show more
tokenmaxxing vs. ironmaxxing lol. it should be an era where results matter but it seems not.