DeepSeek just released its new V4 Flash Vision model, and I plugged it into our humanoid!
This release is exciting because it goes beyond “DeepSeek can now see images.”
DeepSeek is combining the agent/reasoning capabilities of V4 Flash with visual understanding and according to its own benchmarks, the new model already comes close to Anthropic’s Opus 4.8 on multimodal agent tasks.
That immediately made me wonder:
What happens when you give that kind of model the eyes of a real robot? So I gave our robot “RalCox” a simple goal:
“Safely approach the person using the laptop.”
RalCox, sitting besides me, took a fresh image from its onboard camera and sent it to DeepSeek.
In the software screenshot you can actually see the whole experiment: the top-right window is the K1’s live camera view, while the terminal below shows DeepSeek reasoning about what the robot sees.
Instead of asking for a normal image description, I asked it to reason about:
> the current environment
> spatial relationships
>obstacles and action-relevant objects
> what the robot should do next
> what it cannot reliably infer from a single camera image
And that’s where it gets interesting: It identified the table and scattered objects as navigation constraints, reasoned about the likely path toward me, suggested cautiously moving closer while inspecting a potential obstacle and explicitly acknowledged uncertainty around depth, hidden obstacles and surface properties.
Of course, this is not yet a robot that “understands the world.” But to me it feels like a very first building block toward it:
Perception → context → relevance → uncertainty → action reasoning.
And there’s another part I find almost equally fascinating:
I basically don’t code. I started playing with this only a few hours after the model was released and built the whole test with AI helping me step-by-step connecting the robot camera, writing the Python, debugging the API and changing the reasoning prompts.
A few years ago, building even this tiny robotics experiment would probably have required a developer. Today I can have an idea in the afternoon and test it on a humanoid the same evening.
Technically, the setup is still very simple:
Robot camera → Python agent on the robot → DeepSeek V4 Flash Vision → environment reasoning
DeepSeek itself is still running via API, not locally on the robot.
Next experiment: let RalCox actively turn its head, capture several viewpoints, and reason across them before deciding what to do. 🤖
Show more
Today we announce vision capability for our API model. A milestone, and we will continue to evolve. Please try it and share your experience with us!🎉
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!
🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.
Try it now!
Show more
We’re launching DeepSeek-V4-Pro today! 🚀
🔷 Major Agent upgrades with strong production gains!
🔷 Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks.
🔷 Native OpenAI Responses API support, optimized for Codex with one-click setup.
V4 Pro is now available on app/web. Try it via “Expert Mode”.
V4 Pro is also available via API. Model names remain unchanged—please refer to the API docs for setup details.
Show more
Go's first 7T token day
...2 days later
🚀 DeepSeek-V4-Flash Official API is now LIVE in public beta!
🔷 We’ve massively upgraded its Agent capabilities—benchmark scores are now far surpassing the V4-Pro-Preview. Check out the massive performance leap below! 👇
🔷 The official V4-Flash now natively supports the Responses API format and is fully adapted for Codex!
Check out the configuration details in our official API docs:
Show more
Amazing speed and quality.
This #
CVPR2026# paper from our research team is trending #
1# on
@HuggingFace 🤗
Meet LocateAnything: a vision-language detection model that rethinks bounding box prediction. For AI agents and robots, “seeing” is only useful if a model can pinpoint where something is fast enough to act.
Trained on 138M high-quality samples, LocateAnything decodes bounding boxes in parallel instead of one coordinate at a time, improving localization accuracy while dramatically increasing throughput for visual grounding and detection.
Project page:
Show more
We are making our discount permanent! 🎉
Enjoy building with DeepSeek-V4-Pro and bring your innovative ideas to life! 🚀
Fun and honest. It was a joy to read. Science is better when sharing.
We’re talking about Goblins.
The DeepSeek-V4-Pro discount has been extended until May 31, 2026, 15:59 UTC!
DeepSeek v4 Pro is now on Ollama's cloud! 🚀🚀🚀
Try it with Claude Code:
ollama launch claude --model deepseek-v4-pro:cloud
Try it with Hermes Agent:
ollama launch hermes --model deepseek-v4-pro:cloud
Chat with the model:
ollama run deepseek-v4-pro:cloud
🧵
Show more
🚀 Introducing World-R1: Video models already know 3D — they just need RL to wake it up!
No arch changes. No video training data. No extra inference cost.⬇️
🌐Website:
Show more
🔥DeepSeek Input Cache Price Drop!
Effective immediately, the price for input cache hits across the ENTIRE DeepSeek API series is reduced to just 1/10th of the original price! Build more efficiently for less.
📌Reminder: The DeepSeek-V4-Pro 75% OFF promotion is still active until May 5th, 2026, 15:59 (UTC Time).
Show more
🔥DeepSeek-V4-Pro API is 75% OFF until May 5th, 2026, 15:59 (UTC Time)! Don't miss out on this massive discount.
🛠️Integration Updates:
🔹Claude Code: Set model to deepseek-v4-pro[1m] to unlock 1M context!
🔹OpenCode: Update to v1.14.24+
🔹OpenClaw: Update to v2026.4.24+
Check the latest official API docs for full details:
Show more
DeepSeek v4 is now the #
1# open-weight model on our Vibe Code Benchmark, and it’s not close.
It leaves the #
2# (Kimi K2.6) in the dust, and even beats out frontier closed source models like Gemini 3.1 Pro.
Show more
Meet DeepSeek-V4-Pro: our largest and most advanced model to date. Available now on Web, App, and API.
Tech report:
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.
🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.
🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.
Try it now at via Expert Mode / Instant Mode. API is updated & available today!
📄 Tech Report:
🤗 Open Weights:
1/n
Show more
What a year! 2026 will be more exciting.
🚀 Launching DeepSeek-V3.2 & DeepSeek-V3.2-Speciale — Reasoning-first models built for agents!
🔹 DeepSeek-V3.2: Official successor to V3.2-Exp. Now live on App, Web & API.
🔹 DeepSeek-V3.2-Speciale: Pushing the boundaries of reasoning capabilities. API-only for now.
📄 Tech report:
1/n
Show more