Some context first, or the rest won’t make sense: my setup splits work across two machines. One runs Ternary-Bonsai-2-27B through llama-server as a local OpenAI-compatible endpoint; the other connects to it with DeepSeek Harness (DSH below — a model client that lets you plug in your own providers). One is the server, one is the client. Keep that split in mind, because this whole story lives in the gap between them.

My llama-server launch script was pretty loaded:

1
2
3
4
5
6
7
8
9
10
11
12
llama-server.exe ^
-m "%MODEL%" ^
-ngl 99 ^
-c 196608 ^
--flash-attn on ^
--load-mode none ^
--cache-type-k q4_0 ^
--cache-type-v q4_0 ^
--reasoning-preserve ^
--api-key "your-api-key" ^
--host 0.0.0.0 ^
--port 9931

Then, mid-conversation in DSH, this popped up:

Output token limit reached. The response was truncated; what’s been generated so far is kept in the conversation. Send “continue” to let the model pick up where it left off.

My first thoughts: did -c not take effect? Did the KV cache blow up? Is the model just not up to it?

After a round of digging, the answer was blunt: the problem wasn’t llama-server. It was the maxTokens budget on the DSH client side running out.


Why I Ruled Out llama-server First

Output length on llama-server is governed by -n / --n-predict. I didn’t set it, and the default is -1: don’t cap output, just generate until EOS or until the context window is full.

I had -c 196608, a 196k context. If I’d actually hit the window, llama-server would normally complain — context full, OOM, or something in the logs about generation being clipped by the window. There was nothing.

The real tell was that DSH message. Quick background: every OpenAI-compatible response carries a finish_reason, and the two you’ll see most are stop (the model finished on its own) and length (it got cut off by an output cap). DSH’s “send continue” message is its standard fallback when it receives finish_reason: "length".

So the server wasn’t struggling. The output cap in the request simply ran out first.


The Real Trap: contextWindow and maxTokens Didn’t Line Up

Here’s what my DSH config for the local model originally looked like:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
- id: llm-pi-ai
name: "@deepseek-ai/dsh-llm-pi-ai"
config:
providers:
local-llm:
displayName: bonsai-27b
apiKeyEnv: LOCAL_LLM_API_KEY
api: openai-completions
baseURL: http://192.168.3.24:9931/v1
models:
- id: 'D:\UnslothModels\...\Ternary-Bonsai-2-27B-PQ2_0.gguf'
name: Ternary-Bonsai-2-27B-PQ2_0
contextWindow: 32768
maxTokens: 16384
reasoningEffort: none

This is where it went wrong.

Setting Old value What it actually did
contextWindow 32768 DSH thought the window was only 32k
maxTokens 16384 Single response capped at 16384 tokens
llama-server -c 196608 The server could actually run far longer

In other words: the server built a 196k runway, but DSH was only cleared for 16k of output.

Why does an under-reported contextWindow also cause trouble? Because the client doesn’t just send maxTokens and call it a day. It budgets based on how much window is left: the input takes its share, and whatever remains is the output budget. Tell it the window is 32k, and once the input gets long, it starts squeezing well before 16384 — the truncation arrives sooner than you’d expect.

One more thing that’s especially unfriendly to reasoning models: Bonsai is a reasoning model, I had --reasoning-preserve on, and thinking tokens count squarely against the output budget. In the OpenAI ecosystem, reasoning tokens are completion tokens, and local endpoints behave the same way. So the visible reply looked short while the hidden thinking had already burned through thousands of tokens. For a reasoning model, 16384 is much tighter than it looks — that’s fundamentally why reasoning models hit maxTokens more easily than chat models.


One More Sneaky Field: maxTokensField

Raising maxTokens alone isn’t enough. There’s another detail that’s easy to miss.

The OpenAI ecosystem currently has two coexisting field names: max_tokens for the classic API, and max_completion_tokens, which the o1-series reasoning models require. Compatible endpoints have adopted them unevenly. Older llama-server builds only accept max_tokens; newer ones accept both.

The problem: you don’t know which field your client sends, and you may not know which one your llama.cpp build honors. If the two mismatch, you get the spooky outcome of “I raised the limit and nothing changed.”

So when wiring up a local OpenAI-compatible endpoint, pin it explicitly:

1
2
compat:
maxTokensField: max_tokens

One line, and the guesswork is gone. No betting on versions.


My Fixed Config

On the server side, I added one line to the launch script — --alias, to give the model a short name:

1
2
3
4
5
llama-server.exe ^
-m "%MODEL%" ^
--alias bonsai-27b ^
-ngl 99 ^
... (everything else unchanged)

Why bother? Because llama-server defaults to using the model file’s full path as the model id in the API, and a Windows path full of backslashes makes for an ugly, error-prone id. With --alias bonsai-27b, /v1/models returns that short name, and the client config just mirrors it.

Then the DSH yaml became:

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
- id: agent-default-model
name: "@deepseek-ai/dsh-agent-default-model"
config:
provider: local-llm
model: bonsai-27b
- id: llm-pi-ai
name: "@deepseek-ai/dsh-llm-pi-ai"
config:
providers:
local-llm:
displayName: bonsai-27b
apiKeyEnv: LOCAL_LLM_API_KEY
api: openai-completions
baseURL: http://192.168.3.24:9931/v1
compat:
maxTokensField: max_tokens
models:
- id: bonsai-27b
name: Ternary-Bonsai-2-27B-PQ2_0
contextWindow: 196608
maxTokens: 65536
reasoningEffort: none

One trap I actually stepped on, worth its own callout: the model in that first agent-default-model block has to change to bonsai-27b too. I initially only updated the id in the models list below, while the default-model block still pointed at the old path id. The two didn’t match, DSH couldn’t resolve the default model, and my parameter changes silently did nothing. Every place in the yaml that references a model id has to move together.

Two principles:

  • contextWindow must match llama-server’s -c. Otherwise the client budgets against the wrong window and your output budget gets squeezed twice.
  • Don’t max maxTokens out at 196k right away. It’s a per-response cap, and 65536 is roomy enough for most local reasoning tasks. More buys you nothing.

Save and you’re done — the DSH desktop build hot-reloads, no restart needed.


How to Verify the Fix Actually Took

First, confirm which model id DSH is actually hitting:

1
2
curl http://192.168.3.24:9931/v1/models \
-H "Authorization: Bearer your-api-key"

The id in the response must be bonsai-27b, matching the yaml exactly. Note that since my launch script sets --api-key, everything under /v1/* requires auth — curl without the header gets a flat 401. If your client ever starts failing on every request while the llama-server logs look perfectly healthy, check for 401s before blaming the model.

You can also run:

1
dsh --profile desktop --dump-config

to check that agent-default-model and the models entry point at the same id, and that contextWindow and maxTokens show the new values.

Both of those are indirect, though. The decisive move is to fire a real request with a tiny cap and watch it trip:

1
2
3
4
curl http://192.168.3.24:9931/v1/chat/completions \
-H "Authorization: Bearer your-api-key" \
-H "Content-Type: application/json" \
-d '{"model": "bonsai-27b", "messages": [{"role":"user","content":"Count from 1 to 500"}], "max_tokens": 200}'

Check two things in the response: finish_reason should be length, and usage.completion_tokens should stop at exactly 200. If both line up, the max_tokens field is genuinely in effect. This test also demonstrates the “thinking eats budget” point from earlier — a reasoning model often stops mid-count, with barely any visible output, while the completion tokens are already spent.


About “Send Continue”

That message isn’t a bug — it’s a safety net. DSH keeps what’s already generated in the conversation and lets the model pick up from there.

But if you’re running evals, long tasks, or tool chains, I wouldn’t rely on manual continuation. Splicing output by hand tends to introduce duplicated prefixes, and it can corrupt tool-call state. The sturdier fix is to size maxTokens and contextWindow properly so the model finishes in one pass.


Wrap-Up

The biggest takeaway from this one: when you self-host a model, how long the server can run and how long the client lets it run are two separate things.

llama-server’s -c, -n, and KV-cache quantization decide whether it can run;
DSH’s contextWindow, maxTokens, and maxTokensField decide whether it’s allowed to.

Next time you see “output token limit reached,” check three things:

  1. Any context full / OOM in the llama-server logs?
  2. Is the client’s maxTokens too small, and does contextWindow match the server?
  3. Is compat.maxTokensField explicitly set to max_tokens?

That’s local inference for you — layers of parameters stacked on each other, and any layer that doesn’t line up makes it look like “the model isn’t good enough.” Most of the time, the config just isn’t aligned.