Register and share your invite link to earn from video plays and referrals.

Search results for deepseekharness
deepseekharness community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including deepseekharness
🔍 A DeepSeek Harness Plugin That Lets AI Explain the Agent's Own Trace Less than two weeks after DeepSeek Harness entered developer preview, its plugin ecosystem is already taking shape. Zhihu contributor 刘琦 built DSH Trace Insight, a focused plugin for one of the most important but overlooked parts of Agent engineering: understanding what an Agent actually did. 1️⃣ DeepSeek Harness exposes the process Many Agent harnesses behave like black boxes. Users provide a task, wait, and eventually receive an answer. DeepSeek Harness is different. Its trace records tool calls, execution steps, failures, retries, and other intermediate activity. This makes the Agent more transparent, but the raw trace is stored as a large JSONL file. Even with filtering, it is difficult for a human to read and reconstruct the full process. The author's idea was straightforward: If the trace is too complicated for humans, let another AI interpret it. 2️⃣ Turn raw traces into readable analysis DSH Trace Insight adds a side panel that asks an AI model to explain the running trace. It can summarize: 🔹 What the Agent is doing 🔹 Which methods and tools it used 🔹 Where errors or retries occurred 🔹 Whether any risky actions appeared 🔹 What lessons can be extracted from the run Instead of waiting beside an opaque progress indicator, users can inspect how the task is progressing, whether the approach is working, and how risky the current behavior looks. The goal is not to add another capability to the executing Agent. It is to add an interpretability layer around the Agent's behavior. 3️⃣ Use two models for cross-checking When a suspicious step appears, the plugin can send the same trace to two different AI models and compare their analyses. This is useful because trace interpretation is still a model-generated judgment. A second model can expose disagreements, missed risks, or different readings of the same tool call. The plugin can also organize detected issues into a compact list for manual review. That creates a useful three-layer workflow: Agent execution → AI trace analysis → human review It is a lightweight approach to Agent observability without requiring users to inspect thousands of raw log lines. 4️⃣ Installation is intentionally simple The plugin is open source under the MIT license: Users can ask their Agent to install it directly with: Please install this DSH plugin: The current version is designed for the native DeepSeek Harness Web UI. Using it inside third-party desktop wrappers may require additional development, since those clients may package or modify the original Web UI differently. 5️⃣ DeepSeek Harness can become a model worker behind Codex The author also suggests an interesting setup for people who do not use DeepSeek Harness as their primary Agent interface. Open the native DSH Web UI inside Codex's browser. Codex remains the main harness, while it operates DeepSeek Harness and the models connected to it. This creates a layered workflow: 🔹 Codex handles planning and orchestration. 🔹 Lower-cost non-GPT models inside DSH perform lightweight tasks or code inspection. 🔹 DSH Trace Insight exposes how those models executed the work. 🔹 Codex can discuss the results with DSH across multiple rounds, then send the final conclusion to another strong model for an additional review. Compared with assigning every subtask to an expensive model, this setup can reduce cost. Compared with calling another CLI tool blindly, it provides much better visibility into execution. ✅ The real value is observability DSH Trace Insight does one thing: it translates an Agent's raw execution history into something humans can understand. That simplicity is its strength. As Agents begin running longer tasks with more tools and greater autonomy, the important question is no longer just whether they produced the correct answer. We also need to know how they reached it, what failed along the way, and whether they crossed any risky boundaries. 🔗 Full Reading: #DeepSeek# #DeepSeekHarness# #AIAgents# #AgentObservability# #OpenSourceAI# #AIEngineering#
Show more
🧩 DeepSeek and OpenAI Open-Sourced Their Harnesses. The Runtime May Become the Next AI Platform DeepSeek Harness and OpenAI's Codex harness are now open source. But the larger story is not simply that two more repositories became public. Zhihu contributor 第欧根尼 argues that Agent products are beginning to unbundle. The future may be less about choosing one monolithic Agent app and more about assembling a runtime, model router, scenario-specific distribution, and trusted plugin stack. 1️⃣ A harness is becoming part of the model The author's first hypothesis is that a harness will no longer be just a frontend that exposes model capabilities. It will become a framework that co-evolves with the model. The same model may perform very differently inside its official harness than inside a third-party implementation. Context management, tool descriptions, task decomposition, verification, and retry policies all influence the model's behavior. This leads to an important conclusion: A model and its Agent loop can produce better results together than the model can deliver on its own. The competitive unit is therefore shifting from the model alone to the model-harness system. 2️⃣ Agents are becoming lighter and more distributed The author's second observation comes from the evolution of MCP, Cloudflare's Agent infrastructure, and the growing demand for programmable workflows. He expects Agents to become: 🔹 Smaller and more specialized 🔹 Easier to customize through code 🔹 More independent from monolithic apps 🔹 Numerous enough to run as lightweight background workers Current products such as Kimi Work or WorkBuddy still control much of their unique behavior internally. Users cannot easily modify them or embed their complete workflows inside an enterprise system. But market demand is moving toward more flexible forms: plugins lighter than standalone apps, Code Mode more powerful than static skills, and large numbers of low-overhead Agents running simultaneously. That helps explain why vendors are opening their harnesses now. 3️⃣ The harness becomes a microkernel DeepSeek Harness treats the harness as something closer to a microkernel plus a distribution. The base runtime becomes thinner. It retains only the functions every Agent needs: 🔹 Plugin loading and lifecycle management 🔹 Event routing 🔹 Permissions and state 🔹 Execution protocols 🔹 Session and context infrastructure Research, coding, office work, and customer service are then assembled through different plugin bundles. The author sees OpenAI's Codex harness moving in a broadly similar direction, even if it uses different terminology. In this model, users may stop choosing a single Agent product. Instead, they choose: runtime + model routing + scenario distribution + organization plugins DeepSeek Harness and Codex become open runtimes on which many different Agent products can be built. 4️⃣ Five changes follow from this architecture 🔹 Plugin count stops being meaningful Prompts, skills, MCP services, and harness plugins can multiply quickly. The difficult problem will not be finding more plugins, but deciding which ones are trustworthy. Security review, provenance, compatibility, maintenance, and permission control become the real barriers. 🔹 Models become replaceable execution resources If context and data remain inside the harness, the runtime can route different tasks to different models. A strong model may handle planning and review, while cheaper models perform repetitive execution. Switching models becomes a runtime decision rather than a full migration. 🔹 The Agent Loop becomes the main optimization target As model capabilities converge, user experience may depend more on the surrounding loop: When should context be compressed? When should a task be split? What should be remembered? How should results be verified? Improving these decisions may create more value than replacing the underlying model. Models trained to cooperate with a particular harness could gain a significant advantage. 🔹 Skills, MCP, and plugins form a compatibility layer The market is unlikely to accept a different extension format for every platform forever. Competition will shift from “does this platform support plugins?” to “how many ecosystems can it support without degrading the experience?” 🔹 Personal runtimes separate from enterprise control planes Individuals need flexible local Agents. Enterprises need governance, private marketplaces, permission policies, observability, and integration management. These will become distinct product layers, even when they share the same open runtime. 5️⃣ Existing Agent products will defend through ecosystems The author expects products such as WorkBuddy to expose compatibility layers without fully opening their core runtime. They may quickly announce support for DeepSeek Harness plugins, Agent Skills, and more MCP services. But these capabilities would likely enter through adapters rather than replace the underlying harness. They may also build private enterprise plugin marketplaces. The more open the ecosystem becomes, the more companies need vendors that can absorb integration and security risks. Distribution remains another moat. WorkBuddy can connect deeply with WeChat, WeCom, and Tencent Docs. An open harness may reproduce its plugins, but it cannot quickly reproduce users' work relationships and established business entry points. Alibaba has a different advantage. The author expects it to use Alibaba Cloud and Bailian to provide a managed harness control plane, turning cloud infrastructure into the runtime layer for enterprise Agents. 6️⃣ This looks like the Android/AOSP moment for Agents The current market resembles the early Android ecosystem. An open foundation can stop hundreds of teams from rebuilding the same runtime. But publishing reference code is not enough to create an Android-scale platform. The next six months may decide whether these projects converge into a durable ecosystem. They need stable interfaces, trusted plugin infrastructure, and vertical software teams willing to maintain real products on top of the open runtimes. The decisive question is not whether DeepSeek or OpenAI has released the better harness today. It is whether the industry can turn open harnesses into a shared Agent platform, rather than another collection of incompatible reference implementations. 🔗 Full analysis: #DeepSeekHarness# #OpenAI# #Codex# #AIAgents# #AgentInfrastructure# #MCP# #OpenSourceAI#
Show more
DeepSeek Harness v0.1.6-alpha.2 Pre-release with another batch of updates. The app is coming together nicely…
deepseek eval'd their new v4.1 flash model in 8 different harness configs > model performs best in minimal harnesses -- mini-SWE & minimal deepseek harness > underperforms in both claude code and codex
Show more
DeepSeek Harness went from nothing to 150,000 GitHub stars in days, it's MIT licensed, and the plugin system is insane. You can rip out the core, swap the model for Claude or OpenAI, and build your own tooling. Worth a look if Claude Code is your daily driver.
Show more
DeepSeek Harness agents need computers too. Now they can use Sprites.
DeepSeek is having its second DeepSeek moment It's opening up the layer many thought would be the next moat - the harness. Their DeepSeek Harness is the layer that turns a model into an agent: tools, memory, execution loop, sandbox, storage, scheduling, interface. (DeepSeek has released its version of that entire stack under an MIT license) The whole thing is built around one idea: "Everything is a Plugin." Even the model provider. So you don’t have to get the model, tools, and runtime from the same company. ▪️ But here’s the most interesting thing about Harness - the agent can build a missing piece of itself while it’s running. If it needs a tool that doesn’t exist, it can create a temporary plugin. What’s starting to appear is a loop: need a capability → generate the tool → use it → remove it. No retraining or new version. This self-evolving loop isn’t complete yet. But suddenly, you can see the architecture for it. So, R1 challenged the idea that frontier models would stay concentrated in a few labs. DeepSeek Harness is now putting the same pressure on the layer above them. And the next moat suddenly looks a lot more modular and a lot more open.
Show more
The DeepSeek harness is engineering at its finest. 202K stars on GitHub. Damn! Clearly, they built this harness with future AI needs and capabilities in mind. Everything is a plugin and customizable, which is exactly how harnesses of the future need to be.
Show more
With DeepSeek Harness bringing fresh attention to the space, one thing feels clear: the defining test of a harness is whether it can evolve with its developers. That loop is already taking shape around MCode CLI. Developers are using MCode to evolve MCode itself—building everything from a Chinese localization extension to a gloriously unhinged token burner. Their creativity has blown us away. Next week, we’re opening the official MCode CLI Extensions repository and inviting every developer to step into the loop—and help shape what MCode becomes next. Stay tuned.
Show more
🧩 DeepSeek Harness v0.1 is now available in Developer Preview! 🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license. 🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended. Try it now!
Show more
0
844
20.1K
2.5K
Forward to community