llmman launch dsh: Run DeepSeek Harness on any local or hosted model
DeepSeek Harness treats the model as a plugin. llmman runs any model on your own hardware, in one command.
An agent harness is a loop around your model that takes your task, calls a model, runs tools (such as shell commands and file edits), provides results, and repeats. Claude Code, Codex, and OpenCode are all harnesses. DeepSeek Harness is DeepSeek’s.
Unlike Claude Code, DeepSeek Harness is fully customizable. It builds on the idea that everything is a plugin including the model, tools, memory and even the agent loop itself. With DeepSeek Harness, you can build a custom AI agent entirely from scratch, run it locally with any AI model and even; call Claude Code and Codex as sub-agents from inside of it.
Why llmman?
An agent CLI normally talks to one vendor’s API, and your prompts and code go with it. llmman puts a server in between. DeepSeek Harness (dsh) always talks to 127.0.0.1:17434, and you choose what answers: by default a model running on your own machine, or, with --provider, a hosted model from OpenAI, Anthropic, OpenRouter or any of the other providers llmman knows about. Same command, same dsh configuration either way.
Launch DeepSeek Harness with llmman and you get the following out of the box:
- Local by default: Prompts, file contents and diffs stay on your machine. Once the weights are pulled, the loop works offline.
- Hosted when you ask:
--providerchanges where the daemon forwards a request, not whatdshtalks to, so the agent config is the same for local and hosted models, and you can switch between them without touching it. - The model is yours to move:
llmmanstores local models as OCI artifacts. Pull one that is already packaged that way from Docker Hub, or havellmmanpull the weights from Hugging Face and package them for you. Either way, you can then push it to a registry you control, or copy it into a network with no internet at all.
The setup
llmman launch dsh --model <model-name> # eg. gemma4:12b
One command does four things: it starts llmman serve if there isnt a running server, pulls the model and loads it, writes the configuration dsh expects, and hands over to dsh’s web profile. Short names work here the way they do everywhere else in llmman, so gemma4:12b resolves to docker.io/ai/gemma4:12b.
--model is required here, unlike most integrations. dsh has no default model of its own, and leaving it empty writes the literal string default into its settings; which fails at the first request rather than at the command you typed.
Executing a single task
When you run llmman launch dsh --model <model-name>, dsh’s web profile boots a server and answers in a browser, so it has nothing to print to your terminal. When you want one answer and no browser, pass dsh’s headless profile instead. Everything after -- goes to dsh’s own CLI:
llmman launch dsh --model gemma4:12b \
-- --profile headless "Explain what git rebase does in one sentence"
llmman defaults to a web profile. Providing --profile headless overrides the default profile and uses the headless profile.

What llmman configures for you
When you launch dsh with llmman, llmman sets up two files under ~/.config/llmman/launch/dsh. The first registers the daemon as a provider and picks the model:
agent-default-model:
provider: llmman
model: "docker.io/ai/gemma4:12b"
llm-pi-ai:
providers:
llmman:
displayName: llmman
apiKeyEnv: LLMMAN_API_KEY
api: openai-completions
baseURL: "http://127.0.0.1:17434/v1"
models:
- id: "docker.io/ai/gemma4:12b"
name: "docker.io/ai/gemma4:12b"
input: [text, image]
The second is the patch that points dsh at the first. Both are rewritten on every launch, so the model dsh talks to is always the one you just named.
In that configuration:
apiKeyEnvnames an environment variable rather than holding a key, so the credential travels in dsh’s environment and never lands on disk. That is a deliberate security choice: a config file that never holds a credential cannot leak one.inputis populated from local model metadata. llmman asks the daemon what a local model can do, so dsh offers image attachments when a locally served model supports them. Hosted-provider launches currently advertise text input only.
Hosted models
Sometimes you don’t want the model on your machine at all. The weights may not fit on your disk, your hardware may not run them at a useful speed, or the task may need a bigger model than you can run. For those cases, llmman can send dsh’s requests to a model someone else runs.
This is different from pulling a model from Docker Hub or Hugging Face. The registries only store weights; once pulled, the model runs on your machine. A hosted provider runs the model for you, which means your prompts, file contents and diffs go to that provider. That is the “unless you say” from earlier: nothing leaves your machine until you pass --provider.
Run dsh on a hosted provider
Export the provider’s key and add --provider to the same command:
export OPENROUTER_API_KEY=...
llmman launch dsh --provider openrouter --model google/gemma-4-31b-it
The key is read from your environment and sent with each request. It is never written into dsh’s configuration or anywhere else on disk.
Find a provider and a model
The provider list comes from models.dev at runtime, so a provider that appears there works without waiting for an llmman release. There are around 180, including OpenAI, Anthropic, DeepSeek, Groq, Mistral, Together and Fireworks. llmman providers prints each one with the environment variable it expects, whether yours is set, and how many models it serves:
$ llmman providers | grep -E '^(PROVIDER|openrouter)'
PROVIDER NAME API KEY KEY MODELS
openrouter OpenRouter OPENROUTER_API_KEY set 369
To see those models, and the exact name to pass to --model, list them:
llmman list --provider openrouter
Use your own server
A server the catalog has never heard of works too: vLLM on a GPU machine down the hall, LM Studio on a laptop, or a proxy in front of OpenAI. Give it a base_url in llmman.conf and it takes the same flag:
llmman config set providers.local.base_url http://192.168.1.50:8000/v1
llmman launch dsh --provider local --model google/gemma-4-26b-a4b-it
Servers on your own network often need no key. If yours does, set providers.local.api_key_env to the name of the environment variable that holds it. If the server speaks the Anthropic API rather than OpenAI’s, set providers.local.wire to anthropic.
In every case, dsh still talks to llmman serve on 127.0.0.1:17434, with the same generated config. --provider only changes where the daemon sends the request.