Some context first, or the rest won’t make sense: my setup splits work across two machines. One runs Ternary-Bonsai-2-27B through llama-server as a local OpenAI-compatible endpoint; the other connects to it with DeepSeek Harness (DSH below — a model client that lets you plug in your own providers). One is the server, one is the client. Keep that split in mind, because this whole story lives in the gap between them.
My llama-server launch script was pretty loaded:
1 | llama-server.exe ^ |
Then, mid-conversation in DSH, this popped up:
Output token limit reached. The response was truncated; what’s been generated so far is kept in the conversation. Send “continue” to let the model pick up where it left off.
My first thoughts: did -c not take effect? Did the KV cache blow up? Is the model just not up to it?
After a round of digging, the answer was blunt: the problem wasn’t llama-server. It was the maxTokens budget on the DSH client side running out.
Why I Ruled Out llama-server First
Output length on llama-server is governed by -n / --n-predict. I didn’t set it, and the default is -1: don’t cap output, just generate until EOS or until the context window is full.
I had -c 196608, a 196k context. If I’d actually hit the window, llama-server would normally complain — context full, OOM, or something in the logs about generation being clipped by the window. There was nothing.
The real tell was that DSH message. Quick background: every OpenAI-compatible response carries a finish_reason, and the two you’ll see most are stop (the model finished on its own) and length (it got cut off by an output cap). DSH’s “send continue” message is its standard fallback when it receives finish_reason: "length".
So the server wasn’t struggling. The output cap in the request simply ran out first.
The Real Trap: contextWindow and maxTokens Didn’t Line Up
Here’s what my DSH config for the local model originally looked like:
1 | - id: llm-pi-ai |
This is where it went wrong.
| Setting | Old value | What it actually did |
|---|---|---|
contextWindow |
32768 | DSH thought the window was only 32k |
maxTokens |
16384 | Single response capped at 16384 tokens |
llama-server -c |
196608 | The server could actually run far longer |
In other words: the server built a 196k runway, but DSH was only cleared for 16k of output.
Why does an under-reported contextWindow also cause trouble? Because the client doesn’t just send maxTokens and call it a day. It budgets based on how much window is left: the input takes its share, and whatever remains is the output budget. Tell it the window is 32k, and once the input gets long, it starts squeezing well before 16384 — the truncation arrives sooner than you’d expect.
One more thing that’s especially unfriendly to reasoning models: Bonsai is a reasoning model, I had --reasoning-preserve on, and thinking tokens count squarely against the output budget. In the OpenAI ecosystem, reasoning tokens are completion tokens, and local endpoints behave the same way. So the visible reply looked short while the hidden thinking had already burned through thousands of tokens. For a reasoning model, 16384 is much tighter than it looks — that’s fundamentally why reasoning models hit maxTokens more easily than chat models.
One More Sneaky Field: maxTokensField
Raising maxTokens alone isn’t enough. There’s another detail that’s easy to miss.
The OpenAI ecosystem currently has two coexisting field names: max_tokens for the classic API, and max_completion_tokens, which the o1-series reasoning models require. Compatible endpoints have adopted them unevenly. Older llama-server builds only accept max_tokens; newer ones accept both.
The problem: you don’t know which field your client sends, and you may not know which one your llama.cpp build honors. If the two mismatch, you get the spooky outcome of “I raised the limit and nothing changed.”
So when wiring up a local OpenAI-compatible endpoint, pin it explicitly:
1 | compat: |
One line, and the guesswork is gone. No betting on versions.
My Fixed Config
On the server side, I added one line to the launch script — --alias, to give the model a short name:
1 | llama-server.exe ^ |
Why bother? Because llama-server defaults to using the model file’s full path as the model id in the API, and a Windows path full of backslashes makes for an ugly, error-prone id. With --alias bonsai-27b, /v1/models returns that short name, and the client config just mirrors it.
Then the DSH yaml became:
1 | - id: agent-default-model |
One trap I actually stepped on, worth its own callout: the model in that first agent-default-model block has to change to bonsai-27b too. I initially only updated the id in the models list below, while the default-model block still pointed at the old path id. The two didn’t match, DSH couldn’t resolve the default model, and my parameter changes silently did nothing. Every place in the yaml that references a model id has to move together.
Two principles:
contextWindowmust match llama-server’s-c. Otherwise the client budgets against the wrong window and your output budget gets squeezed twice.- Don’t max
maxTokensout at 196k right away. It’s a per-response cap, and 65536 is roomy enough for most local reasoning tasks. More buys you nothing.
Save and you’re done — the DSH desktop build hot-reloads, no restart needed.
How to Verify the Fix Actually Took
First, confirm which model id DSH is actually hitting:
1 | curl http://192.168.3.24:9931/v1/models \ |
The id in the response must be bonsai-27b, matching the yaml exactly. Note that since my launch script sets --api-key, everything under /v1/* requires auth — curl without the header gets a flat 401. If your client ever starts failing on every request while the llama-server logs look perfectly healthy, check for 401s before blaming the model.
You can also run:
1 | dsh --profile desktop --dump-config |
to check that agent-default-model and the models entry point at the same id, and that contextWindow and maxTokens show the new values.
Both of those are indirect, though. The decisive move is to fire a real request with a tiny cap and watch it trip:
1 | curl http://192.168.3.24:9931/v1/chat/completions \ |
Check two things in the response: finish_reason should be length, and usage.completion_tokens should stop at exactly 200. If both line up, the max_tokens field is genuinely in effect. This test also demonstrates the “thinking eats budget” point from earlier — a reasoning model often stops mid-count, with barely any visible output, while the completion tokens are already spent.
About “Send Continue”
That message isn’t a bug — it’s a safety net. DSH keeps what’s already generated in the conversation and lets the model pick up from there.
But if you’re running evals, long tasks, or tool chains, I wouldn’t rely on manual continuation. Splicing output by hand tends to introduce duplicated prefixes, and it can corrupt tool-call state. The sturdier fix is to size maxTokens and contextWindow properly so the model finishes in one pass.
Wrap-Up
The biggest takeaway from this one: when you self-host a model, how long the server can run and how long the client lets it run are two separate things.
llama-server’s -c, -n, and KV-cache quantization decide whether it can run;
DSH’s contextWindow, maxTokens, and maxTokensField decide whether it’s allowed to.
Next time you see “output token limit reached,” check three things:
- Any context full / OOM in the llama-server logs?
- Is the client’s
maxTokenstoo small, and doescontextWindowmatch the server? - Is
compat.maxTokensFieldexplicitly set tomax_tokens?
That’s local inference for you — layers of parameters stacked on each other, and any layer that doesn’t line up makes it look like “the model isn’t good enough.” Most of the time, the config just isn’t aligned.