DeepSeek Harness (DSH) is a coding-agent harness for V4-Pro and V4-Flash, released as npm package @deepseek-ai/dsh on August 13, 2026 — the same day V4-Pro reached general availability. It installs in four steps, runs on a standard DeepSeek API key, and ships alongside a plugin system. One critical thing to know before installing: a completely unrelated Python package called deepseek-harness has been on PyPI since May and also exposes a dsh command. They are not the same tool.
What DeepSeek Harness (DSH) Is
A coding-agent harness manages context selection, tool-error handling, replanning, and completion detection around a model. Benchmark scores shift with the harness config even when the model is fixed, which is why DeepSeek discloses harness settings alongside agent evals rather than reporting raw model numbers.
Familiar examples: Claude Code is a harness for Anthropic's Claude models, Cursor wraps multiple providers into an IDE, and GitHub Copilot is a harness for OpenAI models inside VS Code. An agent is what results when you pair a model with a harness — a system that can edit files, run commands, and call APIs rather than answer single prompts. DSH fills this role for V4-Pro and V4-Flash.
DSH is DeepSeek's first-party harness, built for V4-Pro and V4-Flash the way Claude Code is built for Sonnet and Opus. The July 31 V4-Flash beta changelog entry publicly named it for the first time, stating that agent benchmarks used "DeepSeek Harness minimal mode (to be released soon)" with effort=max, top_p=0.95, and temperature=1.0.
"Minimal mode" is a stripped configuration with maximal reasoning effort and fixed sampling parameters, designed so agent benchmarks are reproducible rather than tuned per-benchmark. The full product — including the plugin system hinted at in beta reports — is broader than the benchmark runner that produced the published scores.
Both V4-Pro and V4-Flash support a 1M-token context window and three reasoning-effort settings: low for simple tasks, high for ordinary agent work, max for complex work (official changelog, August 13 entry).
Installing @deepseek-ai/dsh
Requires Node.js 18 or newer — the same baseline DeepSeek's official Claude Code integration guide specifies.
- Check Node:
node --version(18+ required;brew install nodeon macOS if missing). - Install globally:
npm install -g @deepseek-ai/dsh. - Set your key:
export DEEPSEEK_API_KEY=<your key>from the DeepSeek platform. - Run it:
cd /path/to/my-project && dsh.
macOS permission errors on global npm installs are common. The npm-recommended fix is configuring a user-writable global prefix (npm config set prefix ~/.npm-global); sudo works but can create ownership conflicts with future package updates.
The harness software has no announced separate price. The model API behind it changes on August 16 at 16:00 UTC: V4-family pricing moves to peak/off-peak rates, with off-peak at half of peak (official changelog).
V4-Pro GA Benchmarks — Where the Harness Shows Up
The official numbers DeepSeek published in its August 13 changelog entry, all produced with the harness configuration described above:
| Benchmark | V4-Pro GA | V4-Flash beta |
|---|---|---|
| Terminal Bench 2.1 | 87.9 | 82.7 |
| Cybergym | 83.3 | 76.7 |
| Toolathlon-Verified | 74.1 | 70.3 |
| DSBench-FullStack | 71.1 | 68.7 |
| DSBench-Hard | 67.2 | 59.6 |
| DeepSWE | 62.7 | 54.4 |
| NL2Repo | 61.5 | 54.2 |
| Agents' Last Exam | 25.7 | 25.2 |
| HLE (with tools) | 60.0 | — |
V4-Pro leads Flash by 0.5 to 8.3 points across these benchmarks, with the widest gaps on DeepSWE (+8.3) and NL2Repo (+7.3). Terminal Bench 2.1 at 87.9 is the headline number. An independent public-harness reproduction on r/LocalLLaMA put Flash 0731 at 82.7 across 445 trials, matching DeepSeek's claim within config-variance margins. DSBench-FullStack and DSBench-Hard are internal evaluation sets, so those two rows are DeepSeek's own measurements.
All published benchmarks ran at max effort. For production use, DeepSeek's guidance is low for simple tasks, high for ordinary agent work, and max for complex work.
@deepseek-ai/dsh vs. pip install deepseek-harness
Two packages with near-identical names shipped months apart. They are unrelated.
@deepseek-ai/dsh (npm) | deepseek-harness (PyPI) | |
|---|---|---|
| Maintainer | DeepSeek (official scope) | Henry Zhang, unaffiliated |
| First released | August 13, 2026 | May 9, 2026 |
| Language | Node.js | Python |
| What it is | Coding-agent harness | Protocol-adapter library |
| Install | npm install -g @deepseek-ai/dsh | pip install deepseek-harness |
| License | Not yet verified | MIT |
The PyPI project is a protocol-adapter library, not a coding agent. Its author documented 16 V4 API behaviors across 270+ trials and encoded them as a 10-rule contract shipped in four formats: Python library, dsh CLI, MCP server, and an Anthropic Skill. The three findings most likely to bite anyone building on V4:
- Multi-turn tool loops return HTTP 400 unless
reasoning_contentis preserved on assistant messages — reproduced 3/3 on both models. - The
/betaendpoint silently remapsv4-protodeepseek-reasoner. - Cache hits appear in 256-token blocks after a 1,024-token prefix; one five-turn loop hit 95% equilibrium, cutting effective input cost roughly 50x.
Both packages ship a dsh command, which is the main source of confusion. If you installed via pip, you have the adapter library. If you installed via npm under the @deepseek-ai scope, you have the harness.
Using DeepSeek V4 as a Coding Agent Without DSH
If you want a DeepSeek-powered coding agent before DSH documentation stabilizes, DeepSeek's official integration guide covers three supported paths:
| Tool | Setup | DeepSeek models | Best for | |
|---|---|---|---|---|
| Claude Code | Env vars, Anthropic-compatible endpoint | V4-Pro (main), V4-Flash (subagent) | Terminal coding with max effort | |
| OpenCode | /connect flow, select DeepSeek provider | V4-Pro | Open-source, interactive provider selection (needs v1.14.24+) | |
| OpenClaw | `curl -fsSL https://openclaw.ai/install.sh \ | bash` | V4-Pro or V4-Flash | Personal assistant with Skills, chat integration |
The Claude Code configuration maps DeepSeek models into Claude Code's slots:
export ANTHROPIC_BASE_URL=https://api.deepseek.com/anthropic
export ANTHROPIC_AUTH_TOKEN=<your DeepSeek API key>
export ANTHROPIC_MODEL=deepseek-v4-pro
export ANTHROPIC_DEFAULT_OPUS_MODEL=deepseek-v4-pro
export ANTHROPIC_DEFAULT_SONNET_MODEL=deepseek-v4-pro
export ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v4-flash
export CLAUDE_CODE_SUBAGENT_MODEL=deepseek-v4-flash
export CLAUDE_CODE_EFFORT_LEVEL=max
One mapping quirk: requests for claude-sonnet resolve to Flash, not Pro. DeepSeek's docs route Pro to the Opus slot and Flash to Haiku. All three tools are documented as third-party integrations for which DeepSeek does not guarantee effectiveness or security. For a broader comparison, see our coding-agent LLM comparison and the V4-Pro GA API guide.
FAQ
Is @deepseek-ai/dsh the official DeepSeek Harness?
The @deepseek-ai npm scope matches DeepSeek's GitHub organization, and the release date matches the promised window from the July 31 changelog. Community posts confirm it installs. DeepSeek's August 13 changelog documents V4-Pro GA but does not name the harness, and npm blocks automated verification of the package page.
What's the difference between DSH and deepseek-harness on PyPI?
Different products from different maintainers. @deepseek-ai/dsh is the coding-agent harness on npm, released August 13. pip install deepseek-harness installs an independent Python protocol-adapter library by Henry Zhang (MIT, May 9) that handles V4 API quirks. Both ship a dsh command.
Do I need a DeepSeek API key to use DSH?
Yes. The harness runs on V4-Pro and V4-Flash through the DeepSeek API, so usage bills per token. The npm package itself installs free; the cost is API consumption. V4-family pricing shifts to peak/off-peak rates on August 16 at 16:00 UTC, with off-peak at half of peak.
How does DSH compare to Claude Code?
Claude Code is a mature, documented terminal coding agent that already runs on DeepSeek via the Anthropic-compatible endpoint. DSH's pitch is model-specific optimization: a harness built around V4's protocol behavior rather than adapted to it. That advantage is plausible given the benchmark config disclosure, but unproven in independent use until DSH docs and community testing catch up.
What's the difference between an AI harness and an AI agent?
A harness is the infrastructure layer — context management, tool routing, task orchestration. An agent is the combined system: a model running inside a harness, capable of multi-step actions. Every agent needs a harness, but the same harness can serve different models. Claude Code, Cursor, Copilot, and DSH are all harnesses; what makes them "agents" is the model powering them.
Release Status: What's Confirmed vs. Community-Reported
The facts in this article draw from two tiers of evidence.
Officially confirmed (DeepSeek API changelog, dated entries):
deepseek-v4-proreached GA on August 13 with native OpenAI Responses API support and three reasoning-effort levels (changelog).- V4-Flash entered public beta on July 31, with benchmarks produced using "DeepSeek Harness minimal mode (to be released soon)."
- All V4-Pro GA and V4-Flash beta benchmark numbers come from these changelog entries.
- V4-family pricing changes to peak/off-peak rates on August 16 at 16:00 UTC.
- The Claude Code, OpenCode, and OpenClaw integration guide is official documentation.
Community-reported (Reddit posts, reproduced screenshots):
@deepseek-ai/dshis installable and functional. u/Testx01 posted the package link at 12:56 UTC on August 13; u/Available_Yam_6267 followed within minutes. The@deepseek-ainpm scope matches DeepSeek's GitHub organization.- npm's package page blocks automated fetches, so version, license, and maintainer metadata are unconfirmed through the registry.
"Deepseek harness is out on npm :)" — u/Testx01, r/DeepSeek, August 13
- A reproduced internal-beta message describes a 369-member beta group that pushed a final internal build on August 11 and told plugin developers to tag repositories with
#dshahead of the August 13 public beta, with plugins becoming movable to individual accounts post-launch. - Community reporting cited Chinese tech press in May saying DeepSeek was forming a dedicated harness team and recruiting Agent Harness product roles.
Treat DSH feature, CLI, and plugin claims as community-confirmed until DeepSeek publishes formal documentation on its own site.