Register and share your invite link to earn from video plays and referrals.

StepFun
@StepFun_ai
Scale-up possibilities for everyone. API Platform: HuggingFace:
168 Following    14.6K Followers
We’ve open-sourced onPanda 🐼 — the tool we use internally for LLM data annotation and model inspection. The workflow is simple: find an error, correct the token, and let the model continue. ✍️ Data annotation - 52% lower median annotation time vs. manual post-editing - SFT + preference data in one workflow, with high on-policy fidelity (ΔPPL <1% vs. the model’s resampling baseline) - Precise token-level supervision with paired positive/negative examples, plus agent-trajectory annotation across image, audio, and video 🔎 Model inspection and debugging - Inspect token probabilities and top-k alternatives, steer decoding token by token, and explore SVG generation, web development, and agent tasks directly in the browser. Try it (mobile-friendly): Paper:
Show more
Step 5 Preview pushes our intelligence–cost Pareto frontier outward. It comes in at 44 on the Artificial Analysis Intelligence Index, $0.71 per task. Thanks @ArtificialAnlys for putting Step 5 Preview through the full evaluation. More to come on Oct 15.
Show more
StepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th
Show more
Step 5 Preview works across software environments and sustains execution over long horizons. Its capabilities extend from software engineering and web applications to 3D workflows and programmable hardware. Over longer horizons, Step 5 Preview keeps track of prior results, uses execution feedback to decide what to try next, and continues iterating. We test this behavior in runs lasting up to 24 hours, including tasks involving GPU kernel optimization and automated post-training.
Show more
Introducing Step 5 Preview: Advancing the Pareto Frontier. Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance. - 600B total / 27B active MoE, with 1M context + Vision - Substantially lower task cost at comparable intelligence - Broad software engineering capabilities with sustained execution over long horizons Try Step 5 Preview: Model page: Open weights on Oct 15.
Show more
0
195
1.9K
279
Forward to community
We’re bringing StepAudio 3 to San Francisco this Wednesday. Live demos, an open AMA with the StepAudio team, and a panel with leaders from Cresta, Coval, SGLang-Omni & StepFun on what works, what still breaks in production, and where real-time AI interaction is heading next. 📍 Sep 16 · SF 🎁 $100 API credits + drinks & light dinner Join us ↓
Show more
That’s a wrap on AGNTCon + MCPCon Japan. Two days of meeting AI builders, exchanging ideas on agent infrastructure, and seeing what people are building across the ecosystem. We are especially excited to see several great agent infrastructure projects joining the StepFun Startup Program. Building with AI? We’d love to support what you’re working on: Next stop: San Jose for AGNTCon + MCPCon North America with @AgenticAIFdn (AAIF). More updates from StepFun are coming soon. Stay tuned!
Show more
A great first day at AGNTCon+MCPCon Japan. @AgenticAIFdn (AAIF) We enjoyed meeting builders, exchanging ideas, and sharing what we’ve been working on at StepFun. Missed us today? We’ll be back tomorrow for Day 2. Come find us at Booth T6 and say hi! #AGNTCon# #MCPCon# #AgenticAI#
Show more
Voice AI Night was 🔥 Packed room in Mountain View last night for our keynote on next-gen end-to-end speech models + a founder panel with @PLAUDAI @cresta and Firework on building voice AI at global scale. Thanks Seamate for co-hosting — and thanks to everyone who came out, ate, and asked great questions 🎙️ More builders, more voice AI, soon 👀
Show more
Step Plan Free Trial is now available For a limited time: 🎁 15-day Flash Plan free trial for all users ➕ Unlock another 15 days 🤝 Invite friends to earn up to 90 additional days 🚀 Up to 120 days of free access ⏰ Ends Aug 24 · 23:59 UTC Learn more:
Show more
🎙️ Voice AI Night is coming to Mountain View on July 30. StepFun is teaming up with SEAMATE, PLAUD, Firework, and Cresta to bring together builders across foundation models, AI hardware, video commerce, and enterprise AI. We’ll dive into what it really takes to build Voice AI—from architecture, latency, barge-in, and memory to the challenges of turning great technology into products people actually use. 📍 Mountain View 📅 July 30 🔗 RSVP Link:
Show more
Proud to be part of this collaboration! 🎉 Huge thanks to the vLLM team, Ant Group, and FastAFD. Excited to bring AFD to the open-source community.
🎉 Congrats to the teams behind vLLM AFD Plugin (Ascend & vLLM, @StepFun_ai, @AntGroup, FastAFD): a new experimental plugin under vllm-project that brings Attention-FFN Disaggregation to MoE serving. Attention and the expert/FFN path are two very different workloads that normally share one topology. AFD runs them as separate services, so you can scale attention and experts independently. Same vLLM serving surface, no fork. @NVIDIA GPU and Ascend NPU.
Show more
Open models are becoming a key driver of AI innovation. Excited to see NVIDIA publicly championing this vision. Looking forward to what the next generation of AI builders creates.
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
🟠 StepFun is at #WAIC2026# #StepFun# 🚗 ZEEKR 8X × Super EvaStep into StepFun Amoo World — a treasure-hunt map across 7 stops: 🤖 robots building the Great Wall 🎮 arcade penalty shootout w/ Amoo 📱 STEPX agent phone live demos 🗓️ Jul 17–20 📍 Shanghai World Expo Center · Hall H1 C107 Follow the route. Unlock every stop. 👇
Show more
Great to see Step 3.7 Flash now available through @baseten. We built this model for real-world agent workflows, with native multimodal understanding, reliable tool use and the efficiency needed for production. Thanks to the Baseten team for the great collaboration.
Show more
Step 3.7 Flash is now in the Baseten Model Library! This is a 198B-parameter sparse MoE model (11B active per token) with native image and video input, and a 256K context window. It's a strong option for visual reasoning, agentic coding, and long-context tasks.
Show more
Step 3.7 Flash is currently top 10 on @OpenRouter this month, with 4.29T tokens routed. Builders are pushing it through real agent runs, coding tasks, and long-context workflows. Keep sending the hard stuff.
Show more
Been great seeing Step 3.7 Flash get real use in Nous Portal: people testing, building, and running all kinds of agent workflows with it. We’re keeping free access going with the @NousResearch! Try it out and send us what you make.
Show more
More time to build with Step 3.7 Flash: in partnership with @StepFun_ai, we’re extending the free usage period in Nous Portal by an additional 15 days!
This is the kind of agent workflow Step Plan was built for: connect it once, push through a real build, and keep experimenting without thinking about every single API call. Love the tarot generator demo. Thanks for testing Step 3.7 Flash in Claude Code, @codedailyML 🙌
Show more
I used to dread heavy testing days because every API call felt like watching money disappear in real time. Found a backend built to run flat instead of per call, made specifically for agent work. Pointed my Claude Code setup at it and had it running in under 5 minutes. Here's the full run:
Show more
This is the pain we kept hearing from builders: once an agent starts doing real work, the meter becomes part of the workflow. Step Plan is our attempt to make that less of a distraction. Thanks for putting Step Plan + Step 3.7 Flash through a real Claude Code setup 🙌
Show more
I almost stopped testing new models altogether. Not because they were bad. Because every call left a number climbing in the corner of my screen, and at some point I started watching that number more than I was building. A flat-rate backend fixed that for agent work. I swapped my Claude Code setup to it in under 5 minutes. Here's how:
Show more
Introducing the StepFun Startup Program. We’re supporting early-stage AI teams building real products with StepFun models — from multimodal applications to agentic systems. Selected startups may receive API credits, dedicated ecosystem support, co-marketing opportunities, showcase placement, and warm introductions to selected partners. We’d love to hear what you’re building. Apply now 👇
Show more
Tried Step 3.7 Flash in @cline with Playwright MCP. - Select the Step 3.7 Flash model (free) - Installed Playwright MCP through one-click prompt - Let it open the web page, interact with it, and capture screenshots This shows the workflow: efficient agentic model + good tools + fast browser automation. Step 3.7 Flash is free to try in Cline this month. Cheers guys! 😄
Show more