Register and share your invite link to earn from video plays and referrals.

Xiaokang Chen
@PKUCXK
Researcher in Multimodal Team @deepseek_ai Previously Bachelor & Ph.D at Peking University @PKU1898 Opinions are my own.
63 Following    6.4K Followers
Powered by more extensive RL training, DeepSeek-V4.1-Flash delivers a significant performance boost on visual agent tasks (e.g., agentic visual reasoning) over DeepSeek-V4-Flash-Vision-Exp. 🤓
Welcome to try out this model!😁
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀 🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. 🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model. 1/n
Show more
Nice work!😉
🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now!
Show more
We’re expanding our team! Come and join us! You will be working in Beijing or Hangzhou. Fluent Chinese is required. 各个岗位热招,均可实习,欢迎投递🙌
You can try the following two prompts in the Thinking mode (via web/app) to get a better model experience in certain domains like counting (Note: keep a line break after the bracketed titles):: [Think with Grounding] ....... [Think with Pointing] ...... These two prompts encourage the model to adopt bounding boxes or points (which are classic fundamentals in computer vision) in its thought process. Personally, I love the pointing approach for solving abstract topological/reasoning tasks. Using points to represent continuous trajectories makes the MLLM's reasoning process feel much more human-like. Speaking purely from my personal exploration: Getting a multimodal model to accurately represent continuous trajectories with points is still a highly challenging frontier task for the entire industry. The current performance on real-world scenarios still has a long way to go.
Show more
Vision is now live on web and app. 👀 Come test the new eyes, but give its pure text capabilities a try while you're at it.
Vision is now live on web and app. 👀 Come test the new eyes, but give its pure text capabilities a try while you're at it.
Absolutely mind-blowing work. 🤯 This reminds me of when we were exploring the concept of 'thinking with visual primitives'. We noticed the model excelled in synthetic scenarios, but real-world generalization for complex tasks still faced bottlenecks. Back then, I kept thinking: "If only someone could provide manual annotations for these real-world scenarios..." Using point-form visual primitives for reasoning would be powerful for tasks like state tracking, spatial topology, and logical testing, but they may heavily rely on high-fidelity, real-world labeled data. I think it’s the same across other domains too. The journey toward AGI cannot skip the hard, grounded work of human-labelled data. Massive respect to this effort! 🫡
Show more
@skalskip92 You might not believe it, but I simply manually annotated over 2,000,000 human body parts with ultra-precise detail. Probably no one else could do that.
😎DeepSeek-V4 Preview
🚀 DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length. 🔹 DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models. 🔹 DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice. Try it now at via Expert Mode / Instant Mode. API is updated & available today! 📄 Tech Report: 🤗 Open Weights: 1/n
Show more
With the recent discussions around Gemini 3 Deep Think, I’m revisiting my logs from the Gemini3-Pro-Preview version a few months back. One thing stood out beyond standard benchmarks: its ability to handle scientific long-tail tasks. Specifically, I fed it an image of a complex molecular structure (shown below), and it nailed the formula correctly. It’s an interesting case study on how scientific vision, massive compute, and top-tier research talent converge. Respect to the team. 🫡 #AI# #Gemini# #Science#
Show more
Seeing InternVideo-Next explicitly characterize its architecture as "Encoder–Predictor–Decoder (EPD)" and defining the Predictor as a "latent world model" confirms a clear convergence in the field. 📉 From our early work on Context AutoEncoder (CAE, — where we decoupled the Encoder, Predictor ("Regressor" in the paper), and Decoder — to Meta's I-JEPA/V-JEPA shifting entirely to latent prediction, and now this. It seems we are all validating the same core intuition: Understanding ≠ Pixel Reconstruction. 💡 Decoupling representation learning from high-frequency detail generation to build World Models in latent space is evidently the path forward. Glad to see our early intuition resonating with the latest SOTA. 🫡 #AI# #ComputerVision# #WorldModel# #JEPA# #CAE#
Show more
All about multimodal. 🧐
Demis Hassabis on the next 12 Months: - Full multimodal convergence: Models like Gemini will seamlessly take in and output text, images, audio, and video, with cross-pollination that boosts reasoning + creativity. - Breakthrough visual intelligence: Image models like Nano Banana Pro will produce highly accurate infographics and show near-human visual understanding. - Language + video fusion: Video models integrated with LLMs unlock richer analysis, storytelling, and step-by-step visual reasoning. - World models go mainstream like Genie 3 - Agents become reliable
Show more
Looking at today's LLM / Multimodal LLM benchmarks is like looking at an elite student's exam scores: a high score doesn't guarantee real-world capability. The focus must shift from leaderboards to usability, reliability, and seamless workflow integration.
Show more
Introducing DeepSeek-V3.1: our first step toward the agent era! 🚀 🧠 Hybrid inference: Think & Non-Think — one model, two modes ⚡️ Faster thinking: DeepSeek-V3.1-Think reaches answers in less time vs. DeepSeek-R1-0528 🛠️ Stronger agent skills: Post-training boosts tool use and multi-step agent tasks Try it now — toggle Think/Non-Think via the "DeepThink" button: 1/5
Show more
0
504
14.7K
1.7K
Forward to community
Get Ready for #OpenSourceWeek#!
🚀 Day 0: Warming up for #OpenSourceWeek#! We're a tiny team @deepseek_ai exploring AGI. Starting next week, we'll be open-sourcing 5 repos, sharing our small but sincere progress with full transparency. These humble building blocks in our online service have been documented, deployed and battle-tested in production. As part of the open-source community, we believe that every line shared becomes collective momentum that accelerates the journey. Daily unlocks are coming soon. No ivory towers - just pure garage-energy and community-driven innovation.
Show more
⚠️ Attention The account @XIAOKANG4CHEN is impersonating me. This is a FAKE account and does NOT represent me or my views. Please do not follow or engage with it. #Imposter# #FakeAccount#