I’ve been meaning to do this for a while: run a local LLM on my Windows machine, plug DeepSeek Harness in as the frontend, and make it reachable from other devices on my home network. I ended up with PrismML’s Bonsai 2 27B — a ternary-quantized model whose weights compress down to 6–7GB.
It wasn’t exactly smooth. A few of the pitfalls only show up when you actually try to run the thing. Here’s the full account, in the hope it saves someone else an evening.
Why Bonsai 2
Two things sold me. First, VRAM usage: a consumer GPU with 16GB fits the entire model, no quantization compromises needed. Second, its hybrid attention architecture keeps the KV cache much smaller than a standard model’s, which matters a lot once you push the context length up.
The catch: Bonsai 2 uses non-standard GGUF types (142/143), and only PrismML’s custom llama.cpp fork can load it. My first attempt with stock llama.cpp died immediately with invalid ggml type. Ollama and LM Studio fail the same way. There’s no substitute on the inference-engine side — you need PrismML’s prebuilt binaries.
My Setup
- GPU: RTX 4070 Ti SUPER 16GB, 32GB RAM
- OS: Windows 11, NVIDIA driver 560+
- Inference: PrismML’s prebuilt
llama-server.exe - Frontend: DeepSeek Harness (DSH below)
One thing people agonize over unnecessarily: if your driver ships CUDA 13.4 support, can you run a binary built against CUDA 13.3? Yes. Driver-level CUDA support is backward compatible, and the runtime DLLs bundled in PrismML’s package run fine on newer drivers. I wasted half an hour hesitating over this so you don’t have to.
Where to Put the Model: Nowhere
I downloaded the model through Unsloth Desktop, which buries it in a deep path like:
1 | D:\UnslothModels\hub\models--prism-ml--...\snapshots\...\Ternary-Bonsai-2-27B-PQ2_0.gguf |
My first instinct was to move it somewhere “clean.” I resisted, and that turned out to be right:
llama-servertakes absolute paths —-mpoints straight at the file, wherever it lives;- Unsloth’s directory structure embeds the version hash; moving the file severs that trail;
- GGUF is a universal format. The engine only cares about the path, not whether the folder looks tidy.
The one real rule: no Chinese characters or spaces in the path, or command-line parsing may simply fail.
The Launch Command
Here’s what I ended up running:
1 | llama-server.exe ^ |
The flags that matter:
-ngl 99: offload every layer to the GPU;-c 32768: 32K context — more on how I landed on this below;--flash-attn on: Flash Attention, a noticeable speedup;--load-mode none: disable mmap to avoid page thrashing on Windows;--cache-type-k/v q4_0: quantize the KV cache to Q4 — this is where most of the VRAM savings come from;--host 0.0.0.0: expose the server to the LAN; skip it if you only need local access.
A Real Trap: –load-mode
This one deserves its own section. Older llama.cpp builds used --no-mmap to disable memory mapping. That flag is now deprecated, and the server prints:
1 | DEPRECATED: --mmap and --no-mmap are deprecated. use --load-mode instead |
The valid values are auto / none / mmap / mlock / mmap+mlock / dio, and the modern equivalent of --no-mmap is --load-mode none. I guessed off on my first try and the process exited with an error — no warning, no fallback. Plenty of tutorials still use --no-mmap; copying them verbatim will bite you.
Wiring Up DeepSeek Harness
DSH speaks the OpenAI API format, so configuration is trivial:
- Provider ID:
local-bonsai(any lowercase identifier) - Display name:
Bonsai 27B - API URL:
http://127.0.0.1:8080/v1— the trailing/v1matters; without it, nothing connects - Protocol:
OpenAI Chat Completions - API key:
abc-12345— anything that matches your launch command; a placeholder is fine locally
Once saved, DSH’s extras — web search, agents — all work on top of the local model.
Opening It to the LAN
On the server, --host 0.0.0.0 is already in the launch command. The only remaining step is a firewall rule for port 8080 (run once in an elevated PowerShell):
1 | netsh advfirewall firewall add rule name="llama-server-8080" dir=in action=allow protocol=TCP localport=8080 |
On the client side, swap 127.0.0.1 in the DSH API URL for the server’s LAN IP, e.g. http://192.168.x.x:8080/v1.
A crude but reliable connectivity check: from another device, open http://<server-ip>:8080 in a browser. If llama-server’s built-in page loads, the path is clear.
How Much Context Can It Actually Take?
Bonsai 2 advertises a 262K context ceiling. On 16GB of VRAM, here’s what I measured:
| Context | Result |
|---|---|
| 32K | Stable, normal speed |
| 64K | Stable, slightly slower |
| 128K | No crash, but VRAM nearly maxed |
| 160K+ | Works, but time-to-first-token gets painful |
My test method was deliberately crude: change only -c, restart, send a fixed prompt, and watch three things — whether dedicated VRAM in Task Manager pins above 15.5GB, whether first-token latency stretches from seconds to tens of seconds, and whether generation speed falls off a cliff.
My takeaway: 32K–64K is the daily sweet spot; 128K is for the occasional long document; beyond 160K it technically runs but you won’t enjoy waiting for it.
Pitfall Cheat Sheet
Everything that went wrong, in one place — mostly for future me:
| Symptom | Cause | Fix |
|---|---|---|
invalid ggml type 142 |
Non-PrismML llama.cpp | Use PrismML’s prebuilt binaries |
--load-mode: invalid value |
Illegal value like off |
Use none |
| LAN clients can’t connect | Firewall | Allow port 8080 |
| Abnormally slow generation | Windows WDDM memory management | Add --load-mode none, lower -c if needed |
| Port already in use at launch | Something else on 8080 | Switch to --port 8081 or similar |
One-Click Launcher
Finally, I froze the whole thing into a bat file sitting next to llama-server.exe — double-click and it runs:
1 | @echo off |
The only line you need to touch is MODEL — point it at your actual path.
Closing Thoughts
Looking back, the whole setup is almost boring: a custom engine, an absolute path, an OpenAI-compatible frontend, and one firewall rule. No Docker, no Python environment — double-click a bat file and you’re running. For people who just want to quietly use a local model, “unremarkable but actually works” might be exactly what’s needed.
The practical value of a local model is real to me: no token burn, nothing leaves the machine, and it works offline. The fact that 16GB of VRAM can carry a 27B model at all felt worth writing down.
I’m an indie developer currently going through the New Zealand skilled migration process. This blog documents my vibe coding experiments, one-person-company lessons, and migration journey.