Live failure: deepseek-v4-flash answered a greeting entirely in the
reasoning channel — no text block, no tool calls. The serializer's
null-content fallback produced an assistant message with neither
content nor tool_calls, which the API 400s ('Invalid assistant message:
content or tool_calls must be set'). Because that message sits durably
in the session log, every later turn of the session re-derived the same
history and failed identically — one all-reasoning response bricked the
session permanently (log: turns 2 and 3 failing byte-identically).
content is now always the flattened text ('' when there is none); the
passback rule still keeps reasoning_content off plain turns. The old
null shape was pinned by a test whose comment claimed the wire accepts
it — live-falsified, updated together with the code, plus a regression
test for the reasoning-only shape. Existing bricked logs resume cleanly
under the fix (the poisoned message now serializes as '').
@deepseek-ai/dsh-llm-deepseek
DeepSeek chat-completions adapter for the harness LLM seam: hand-rolled fetch + SSE translation from the official wire format (source of truth: the API docs — guides/thinking_mode, guides/tool_calls, api/create-chat-completion) into the StreamChunk protocol.
A second, independent implementation of the same seam exists in @deepseek-ai/dsh-llm-pi-ai (library-backed). Same Config shape — pick one per context (registering both for the same model names throws by design).
Config
- id: llm-deepseek
name: '@deepseek-ai/dsh-llm-deepseek'
config:
apiKey: !!js process.env.DEEPSEEK_API_KEY # or rely on the env fallback
baseURL: !!js process.env.DEEPSEEK_BASE_URL # default: https://api.deepseek.com
models: [deepseek-v4-flash, deepseek-v4-pro] # one adapter, registered for each name
thinking: enabled # optional; provider default is enabled
reasoningEffort: high # optional; high | max — omitted ⇒ not sent
models lists every model name this one adapter instance serves: the adapter registers itself for each (the harness model name IS the wire model string), so a generate/stream call routes to it whenever options.model is any of them. Registering a second adapter for a name already taken throws LlmError('DUPLICATE_ADAPTER') (the LLM service enforces one adapter per model, all-or-nothing).
reasoningEffort is omitted by default — when unset, the reasoning_effort wire field is not sent and the server applies its own default for the model. The only accepted values are high and max (DeepSeek's official effort levels). It is meaningful only with thinking enabled (the provider default).
thinking/reasoningEffort are adapter-level request defaults serialized as the official top-level thinking: {type} / reasoning_effort wire fields. They live in adapter config (not GenerateOptions) to keep the core vocabulary provider-neutral.
App attribution
Every request carries the shared attribution header from dsh-llm's attributionHeaders() - the mandatory User-Agent baseline identifying the harness (see dsh-llm § App attribution). Direct DeepSeek requests and OpenAI-compatible gateway requests get no provider-specific app-attribution headers under this adapter contract; OpenRouter app attribution is deferred to a future explicit OpenRouter adapter or mode.
Wire-format notes (verified live + against the official docs)
- Streaming only (
stream_options.include_usagealways on).usagemay arrive attached to the finish chunk or as a trailing usage-only chunk — the translator defers both to[DONE], sousagealways precedesfinishand nothing followsfinish. - The first thinking-mode chunk carries
reasoning_content: ""— handled (no spurious reasoning block). - Reasoning passback rule: on assistant turns that carried tool calls,
reasoning_contentis serialized back in history (required by the API in thinking mode); on tool-call-free turns it is dropped (ignored anyway — saves tokens). - Cache accounting:
cacheReadTokens←prompt_cache_hit_tokens/prompt_tokens_details.cached_tokens; DeepSeek reports no cache-write metric.
Limitations (MVP, documented deliberately)
tool_choiceis not mapped (not part of the core vocabulary).
Errors
Non-2xx responses throw LlmError with stable codes: AUTH (401/403), RATE_LIMIT (429), INVALID_REQUEST (400), SERVER (5xx), HTTP_<status> otherwise. Protocol violations throw STREAM_CLOSED (no [DONE]) or MALFORMED_RESPONSE (bad JSON payload). Unknown wire finish_reasons (e.g. content_filter, insufficient_system_resource) become finish {kind: 'error', code: <REASON>} chunks.
Testing
Unit suites run against a local node:http mock SSE server (no network). Real-API coverage lives in tests/adapter.e2e.ts (pnpm run test:e2e, key-gated): V4 Flash + V4 Pro across thinking enabled/disabled and both official effort levels, including the thinking+tools round trip with reasoning passback.