An htop-style overview of local LLM backends — llama.cpp, Ollama and Lemonade Server — showing which model is loaded, what it costs in memory and what is going through it right now, plus integrated GPU and NPU state.
Drawn like btop: framed panels with titles set into the border, braille history graphs with a colour gradient, a clock, and a background of its own. One file, Python 3.11+, no dependencies.
╭─┐tower · GMKtec NucBox EVO-X2┌─────────────────┐11:19:05┌──────────────────────────────────────┐up 1d18h┌─╮
│ AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 32 th… │ Radeon 8060S · 2737Mhz │
│ CPU ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀ 9% │ GPU ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣿⣿⣿⣿⣿⣿⣿⣿⣿ 100% │
│ ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀ │ ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣿⣿⣿⣿⣿⣿⣿⣿⣿ │
│ RAM 46.6 GiB/124.9 GiB 37% │ GTT 40.3 GiB/120.0 GiB 63°C 85W 34% │
│ ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣤⣤⣤⣤⣤⣤⣤⣤⣤⣤ │ ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣤⣤⣤⣤⣤⣤⣤⣤⣤⣤ │
│ load 1.93 2.27 1.65 │ NPU idle amdxdna · fw 1.1.2.65 · D3hot │
╰─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────╯
╭─┐llama.cpp┌─────────────────────────────────────────┬──────────────────────────────┐2 asleep · 1 stopped┌─╮
│ ○ llama-qwen asleep :8091 socket-activ… │ │
│ qwen3-30b-a3b · fa · ngl 999 │ - ctx 32k slots - idle 2h18 │
│ ○ llama-qwen36 asleep :8090 socket-activ… │ │
│ qwen3.6-35b-a3b · draft draft-mtp · fa · ngl 999 │ - ctx 64k slots - idle 23h15 │
│ ○ llama-v4 stopped dead │ │
│ deepseek-v4-flash · ngl 999 · model file missing │ - ctx 64k slots - │
╰─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────╯
╭─┐Ollama┌────────────────────────────────────────────┬────────────────────────────────┐1 busy · 1 running┌─╮
│ ● ollama running busy :11434 │ cpu 2.0 cores gpu - 23.7 tok/s ⣀⣸⣿⣿⣿⣿⣿⣿⣿⣿ │
│ granite4.2:latest · v0.34.1-igpu-trust-vram │ 25.1 GiB GPU ctx 128k slots 1/1 up 58s │
╰─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────╯
╭─┐Lemonade┌──────────────────────────────────────────┬────────────────────────────────┐1 busy · 1 running┌─╮
│ ● lemonade running busy :13311 front: Gemma… │ cpu 0.0 cores gpu - - tok/s │
│ 2 models · v11.6.0 · 188.1k tok total │ 959.6 MiB RSS ctx - slots - up 20h56 │
│ ● SDXL-Turbo running :8002 ready │ cpu 0.0 cores gpu - - tok/s │
│ SDXL-Turbo · sd-cpp/gpu · image │ 21.3 MiB RSS ctx 32k slots - used 2h34 a… │
│ ● Gemma-4-E4B… running busy :8001 │ cpu 0.3 cores gpu - 44.5 tok/s ⣀⣿⣿⣿⣿⣿⣿⣿⣿⣿ │
│ Gemma-4-E4B-it-GGUF · llamacpp/gpu │ 938.2 MiB RSS ctx 128k slots 2/4 │
╰─┘q quit +/- interval r refresh└───────────────────┴───────────────────────────┘llmtop 0.9.0 · every 1s└─╯
btop has no extension point — shown_boxes accepts only cpu mem net proc and
gpu0…gpu5, and the boxes are hard-wired in its C++ source. A box of your own
means a fork that needs rebasing on every release. So llmtop runs next to btop
rather than inside it, and borrows its look: braille graphs, gradient colours,
meters that fill left to right.
For Ollama alone there are already good tools — otop, ollama-tui, OllamaManager. None of them knows about llama.cpp or Lemonade, and none is prepared for socket activation.
This is why a generic tool is not enough here. When a llama.cpp backend hangs off
a systemd socket unit, an HTTP status check is itself enough to load the
model — twenty seconds and twenty gigabytes to answer "is this running?".
llmtop never talks to a socket port. State comes from systemd, and measurements only from the internal backend port, and only while the service is already up. The socket ports go on a block list before the first HTTP call is made.
Model, context size and options of a sleeping backend are read from the unit:
ExecStart via systemctl show (where %h and friends are already expanded),
and if that points at a start script, the script is parsed — including the case
where its arguments are collected in a bash array first. If the model file has
since disappeared, it says so.
git clone https://github.com/huppiflupp/llmtop.git
install -m 755 llmtop/llmtop.py ~/.local/bin/llmtop
llmtopllmtop # TUI with live graphs
llmtop -n 1 # refresh every second
llmtop --once # print once and exit (meters instead of history)
llmtop --json # machine readable, for scripts and status bars
llmtop --graph-height 3 # taller graphs; default adapts to the window
llmtop --background 16 # another xterm-256 background, or "none"
llmtop --theme orange # a colour theme: default, orange, or any btop theme
llmtop --ascii # no braille or box drawing, plain ASCII
Keys: q quit, m menu, +/- interval, r refresh now.
m opens a menu over the panels, as in btop: the settings on one tab, the
endpoints on the other (Tab switches). ↑/↓ pick a setting, ←/→ change
it, and every change applies at once - theme, graph colours, interval, graph
height, background, ASCII mode. On the endpoints tab a adds a server (type
its URL, Enter), u and n edit URL and name, d deletes, and Enter
tests the selected one right away: the line below shows what answers there
(llama.cpp · b392 · qwen3.6-35b-a3b · ctx 256k) or unreachable. New
endpoints appear in the Endpoints panel immediately. s writes everything to
the config file, Esc closes the menu.
Saving edits the file line by line: only the keys the menu knows are touched,
comments and other sections stay, and the previous file is kept as
config.toml.bak.
theme in [ui] (or --theme) picks the colours: default is llmtop's own,
orange is bundled, and any btop theme works by name (dracula, nord,
gruvbox_dark, …) when btop is installed - llmtop looks in
~/.config/llmtop/themes, ~/.config/btop/themes and /usr/share/btop/themes -
or by path to a .theme file. btop's keys are mapped onto llmtop's styles:
text, titles and highlights as in btop, the four box colours onto the panel
frames, div_line onto the empty part of the graphs, the CPU gradient onto
the graphs and meters, and temp_end onto "busy". Colours are quantised to
xterm-256, so a theme looks close to btop's, not identical.
graph_colors = "default" keeps llmtop's green-yellow-red gradient inside a
theme's frames and text - orange with colourful graphs. background = "theme"
(the default) takes the theme's background, "none" the terminal's own.
Each panel spans the full width and, from 90 columns on, is split down the middle: identity and system on the left, GPU and live measurements on the right. Every field has a fixed column - values are padded, never shifted - so nothing wanders as numbers change length. Below 90 columns the two halves are stacked inside the same panels.
Graphs and meters grow and shrink with the window, trailing fields are cut with an ellipsis rather than pushed aside, and the graphs take up whatever rows the panels leave free, so the screen fills at any window height. As in btop, load and memory share that space: CPU with RAM, GPU with GTT. On short terminals the graphs flatten to one row before the bottom panels would get cut off. Resizing keeps the history: the sample buffer is far wider than any terminal, so a wider window simply reveals more of the past.
The top border carries host and machine name, the clock and uptime; the CPU and
GPU names head their halves of the system panel. Each backend panel lists a
summary such as 1 busy · 2 asleep in its border, and the bottom border carries
the key bindings.
| Reading | Source |
|---|---|
| llama.cpp: state, model, context | systemctl show on service and socket, start script |
| llama.cpp: slots, tok/s | GET /slots on the internal port, delta of n_decoded |
Endpoints ([endpoints] urls) |
name: host in the URL, reverse lookup of an address, or the configured name; engine: owned_by in GET /v1/models, confirmed once by /props, /api/version, /version or /api/v1/health; model: /v1/models, for Ollama /api/ps, for Lemonade /api/v1/health; /slots if it has one, else /metrics: busy from requests_processing, tok/s from tokens_predicted_total ÷ tokens_predicted_seconds_total of finished requests |
| Ollama: model, memory, unload timer | GET /api/ps (size_vram, context_length, expires_at) |
| Ollama: model ↔ runner process | manifests under models/manifests, blob digest of the model layer |
| Lemonade: models, backends, idle time | GET /api/v1/health, last_use against /proc/uptime |
| Lemonade: throughput | GET /api/v1/stats, else /slots of the matching llama.cpp |
| Memory per process | /proc/<pid>/fdinfo, drm-resident-gtt + drm-resident-vram |
| GPU time per process | /proc/<pid>/fdinfo, delta of drm-engine-* |
| GPU overall | /sys/class/drm/card*/device, else nvidia-smi |
| NPU | /sys/class/accel/*, xrt-smi examine, open /dev/accel/* handles |
| Benchmarks and measurement runs | processes by program name (llama-bench, llama-perplexity, colibri, …), /proc/<pid> and its fdinfo |
Programs that load the machine without serving an API, such as llama-bench,
llama-perplexity or a colibri run, get their own benchmarks panel. They
have no tok/s to read, but without the panel their load showed up with no
source at all. The list of program names is configurable ([bench]). A binary
of another name inside a folder of a listed name counts too, so a colibri run
built per model (~/src/colibri-27b/c/qwen36) is still a colibri row, with
its command line on the line below.
Process CPU is given in cores (1.0 = one full core), so it cannot be mistaken for the whole-machine CPU graph above it. The device a backend actually runs on is marked in bold red: CPU when it uses a larger share of all logical CPUs than of the GPU, GPU the other way round, nothing below 10% of either.
A service running as its own user (Ollama, Lemonade) hides its fdinfo, so its
own GPU time cannot be read. When exactly one backend is working, it is given
the device load that the readable processes leave over, shown as ~90%; with
two or more working at once llmtop does not guess and prints -.
tokens_predicted_total from /metrics is only a fallback: that counter is
written when a task finishes and stands still during generation. /slots counts
along live.
A server that is neither a llama-* unit nor a llama-server process — one in a
container, on another host, or an engine with its own binary name — can be listed
under [endpoints] urls, as a URL or as {url = "…", name = "…"}. Such servers get
their own Endpoints panel and are measured over HTTP only, so it shows no process
CPU, memory or GPU figures. The row is named after the host (ai395:8090: the name
in the URL, a reverse lookup of the address, or the configured name), and the detail
field says which program answers — llama.cpp, Ollama, vLLM, LM Studio or Lemonade,
told apart by owned_by in /v1/models and confirmed by a version probe (/props,
/api/version, /version, /api/v1/health) once per server; the version goes on the
line below. An Ollama or Lemonade endpoint shows the model it has loaded (from
/api/ps or /api/v1/health), not the first of its library; a local Ollama already
shown in its own panel is skipped. Without
/slots, its tok/s is the speed of the last finished request, taken from the
two /metrics counters rather than as a rate over wall time (which would jump
when a request ends and fall back to zero). It is shown as tok/s only while the
server is busy; when it is idle the column shows a dimmed 0.0 and the figure moves to the
line below (last request … tok/s), so an idle server does not look busy.
On unified-memory systems (AMD Strix Halo and relatives) the model lives in GTT
and does not appear in RSS at all — a 30B model shows up there as 99 MiB. Only
drm-resident-gtt reveals the real 18.5 GiB. llmtop reports both, separately.
- NPU utilisation in percent is only exposed by
amdxdnathrough debugfs, which needs root. Without it llmtop shows whether anything holds/dev/accel/*open, plus driver, firmware version and power state. - Per-process memory and GPU time are only readable for your own processes.
If Ollama runs as its own system user, what remains for its runners is the
figure from
/api/ps— good, but coarser. - With several Ollama models loaded at once, matching model to process works only when the manifests are readable.
Without any configuration the usual addresses are tried and services are found by
their processes; OLLAMA_HOST, OLLAMA_MODELS and LEMONADE_URL are honoured.
For anything unusual: ~/.config/llmtop/config.toml, see
config.toml.example. The [ui] settings and the
endpoints can also be changed and saved from the menu (m).
MIT

