Serving your first model#
Prove the install with a small one#
Open http://127.0.0.1:8080. The panel shows one list of every model, with the ones already on this disk at the top under on disk and the rest under available. Find Qwen3-8B — 5 GB, so you find out the download and serve path works without waiting for 15 GB — and press Get.
Progress streams under its row. When it finishes the row moves up into on disk and its button becomes Start; press that.
Clicking a row’s name opens what it is and what it is for, which is worth reading before you spend 17 GB on one of the larger ones.
Watch the engine log at the bottom of the panel. A first load looks like:
load_model: loading model '/home/giles/models/Qwen3-8B/Qwen3-8B-Q4_K_M.gguf'
...
srv load_model: the model is loaded
The panel’s engine pill goes stopped → loading → running. The distinction
matters: llama-server binds its HTTP port long before the weights finish
uploading to the card, so “the port is open” is not “the model is ready”.
Talk to it#
curl -s http://127.0.0.1:1919/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Qwen3-8B","messages":[{"role":"user","content":"Hello"}],"max_tokens":30}'
Or from the CLI:
lllm3090 status
Note
Qwen3-8B is a reasoning model and thinks at length before answering. With a
small max_tokens you may get a thinking block and no text at all — that is
the model spending its budget, not a broken endpoint. Give it a few hundred
tokens of headroom.
Move to a real model#
Qwen3-8B is a smoke test — and too small for an agent harness, since its 32k
window cannot hold Claude Code’s ~40k system prompt. For actual work:
Qwen3.6-35B-A3B (17.7 GB) is the default recommendation. Measured at
126 tok/s, and its hybrid attention makes the KV cache cheap enough
(20 KiB/token) to reach the model’s full 262k context — 212k per
conversation with a slot spare for a subagent.
gpt-oss-20b (12.1 GB) is faster still at 160 tok/s and leaves most of
the card free, at 128k of context across four slots. Worth having.
Qwen3.8-27B (15.4 GB) is the dense option. It is the slowest model here
at 35 tok/s — every one of its 27B parameters is read per token, where the
sparse models read about 3B — and it gives 101k per conversation. Take it if
you want a dense model’s qualities and can afford a quarter of the speed; the
comparison is Dense and sparse models want different hardware.
Free the card#
lllm3090 stop
or press Stop in the panel. Do this before gaming, or before anything else that wants the VRAM — a loaded model holds 12–21 GB indefinitely.