Skip to content

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

llmtop

An htop-style overview of local LLM backends — llama.cpp, Ollama and Lemonade Server — showing which model is loaded, what it costs in memory and what is going through it right now, plus integrated GPU and NPU state.

Drawn like btop: framed panels with titles set into the border, braille history graphs with a colour gradient, a clock, and a background of its own. One file, Python 3.11+, no dependencies.

llmtop with the orange theme and the default graph colours, a llama-server busy at 156 tok/s

╭─┐tower · GMKtec NucBox EVO-X2┌─────────────────┐11:19:05┌──────────────────────────────────────┐up 1d18h┌─╮
│      AMD RYZEN AI MAX+ 395 w/ Radeon 8060S · 32 th… │      Radeon 8060S · 2737Mhz                         │
│ CPU  ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀    9% │ GPU  ⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⠀⣿⣿⣿⣿⣿⣿⣿⣿⣿  100% │
│      ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀       │      ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣿⣿⣿⣿⣿⣿⣿⣿⣿       │
│ RAM  46.6 GiB/124.9 GiB                         37% │ GTT  40.3 GiB/120.0 GiB            63°C   85W   34% │
│      ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣤⣤⣤⣤⣤⣤⣤⣤⣤⣤       │      ⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣀⣤⣤⣤⣤⣤⣤⣤⣤⣤⣤       │
│ load 1.93  2.27  1.65                               │ NPU  idle          amdxdna · fw 1.1.2.65 · D3hot    │
╰─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────╯
╭─┐llama.cpp┌─────────────────────────────────────────┬──────────────────────────────┐2 asleep · 1 stopped┌─╮
│ ○ llama-qwen     asleep        :8091  socket-activ… │                                                     │
│   qwen3-30b-a3b · fa · ngl 999                      │         -      ctx   32k  slots     -  idle 2h18    │
│ ○ llama-qwen36   asleep        :8090  socket-activ… │                                                     │
│   qwen3.6-35b-a3b · draft draft-mtp · fa · ngl 999  │         -      ctx   64k  slots     -  idle 23h15   │
│ ○ llama-v4       stopped              dead          │                                                     │
│   deepseek-v4-flash · ngl 999 · model file missing  │         -      ctx   64k  slots     -               │
╰─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────╯
╭─┐Ollama┌────────────────────────────────────────────┬────────────────────────────────┐1 busy · 1 running┌─╮
│ ● ollama         running  busy :11434               │ cpu  2.0 cores  gpu      -    23.7 tok/s ⣀⣸⣿⣿⣿⣿⣿⣿⣿⣿ │
│   granite4.2:latest · v0.34.1-igpu-trust-vram       │  25.1 GiB GPU  ctx  128k  slots   1/1  up 58s       │
╰─────────────────────────────────────────────────────┴─────────────────────────────────────────────────────╯
╭─┐Lemonade┌──────────────────────────────────────────┬────────────────────────────────┐1 busy · 1 running┌─╮
│ ● lemonade       running  busy :13311 front: Gemma… │ cpu  0.0 cores  gpu      -       - tok/s            │
│   2 models · v11.6.0 · 188.1k tok total             │ 959.6 MiB RSS  ctx     -  slots     -  up 20h56     │
│   ● SDXL-Turbo   running       :8002  ready         │ cpu  0.0 cores  gpu      -       - tok/s            │
│     SDXL-Turbo · sd-cpp/gpu · image                 │  21.3 MiB RSS  ctx   32k  slots     -  used 2h34 a… │
│   ● Gemma-4-E4B… running  busy :8001                │ cpu  0.3 cores  gpu      -    44.5 tok/s ⣀⣿⣿⣿⣿⣿⣿⣿⣿⣿ │
│     Gemma-4-E4B-it-GGUF · llamacpp/gpu              │ 938.2 MiB RSS  ctx  128k  slots   2/4               │
╰─┘q quit  +/- interval  r refresh└───────────────────┴───────────────────────────┘llmtop 0.9.0 · every 1s└─╯

Why not just extend btop

btop has no extension point — shown_boxes accepts only cpu mem net proc and gpu0…gpu5, and the boxes are hard-wired in its C++ source. A box of your own means a fork that needs rebasing on every release. So llmtop runs next to btop rather than inside it, and borrows its look: braille graphs, gradient colours, meters that fill left to right.

For Ollama alone there are already good tools — otop, ollama-tui, OllamaManager. None of them knows about llama.cpp or Lemonade, and none is prepared for socket activation.

Socket-activated backends

This is why a generic tool is not enough here. When a llama.cpp backend hangs off a systemd socket unit, an HTTP status check is itself enough to load the model — twenty seconds and twenty gigabytes to answer "is this running?".

llmtop never talks to a socket port. State comes from systemd, and measurements only from the internal backend port, and only while the service is already up. The socket ports go on a block list before the first HTTP call is made.

Model, context size and options of a sleeping backend are read from the unit: ExecStart via systemctl show (where %h and friends are already expanded), and if that points at a start script, the script is parsed — including the case where its arguments are collected in a bash array first. If the model file has since disappeared, it says so.

Install

git clone https://github.com/huppiflupp/llmtop.git
install -m 755 llmtop/llmtop.py ~/.local/bin/llmtop
llmtop

Usage

llmtop                  # TUI with live graphs
llmtop -n 1             # refresh every second
llmtop --once           # print once and exit (meters instead of history)
llmtop --json           # machine readable, for scripts and status bars
llmtop --graph-height 3 # taller graphs; default adapts to the window
llmtop --background 16  # another xterm-256 background, or "none"
llmtop --theme orange   # a colour theme: default, orange, or any btop theme
llmtop --ascii          # no braille or box drawing, plain ASCII

Keys: q quit, m menu, +/- interval, r refresh now.

Menu

m opens a menu over the panels, as in btop: the settings on one tab, the endpoints on the other (Tab switches). ↑/↓ pick a setting, ←/→ change it, and every change applies at once - theme, graph colours, interval, graph height, background, ASCII mode. On the endpoints tab a adds a server (type its URL, Enter), u and n edit URL and name, d deletes, and Enter tests the selected one right away: the line below shows what answers there (llama.cpp · b392 · qwen3.6-35b-a3b · ctx 256k) or unreachable. New endpoints appear in the Endpoints panel immediately. s writes everything to the config file, Esc closes the menu.

The menu over the panels

Saving edits the file line by line: only the keys the menu knows are touched, comments and other sections stay, and the previous file is kept as config.toml.bak.

Themes

theme in [ui] (or --theme) picks the colours: default is llmtop's own, orange is bundled, and any btop theme works by name (dracula, nord, gruvbox_dark, …) when btop is installed - llmtop looks in ~/.config/llmtop/themes, ~/.config/btop/themes and /usr/share/btop/themes - or by path to a .theme file. btop's keys are mapped onto llmtop's styles: text, titles and highlights as in btop, the four box colours onto the panel frames, div_line onto the empty part of the graphs, the CPU gradient onto the graphs and meters, and temp_end onto "busy". Colours are quantised to xterm-256, so a theme looks close to btop's, not identical.

graph_colors = "default" keeps llmtop's green-yellow-red gradient inside a theme's frames and text - orange with colourful graphs. background = "theme" (the default) takes the theme's background, "none" the terminal's own.

Layout

Each panel spans the full width and, from 90 columns on, is split down the middle: identity and system on the left, GPU and live measurements on the right. Every field has a fixed column - values are padded, never shifted - so nothing wanders as numbers change length. Below 90 columns the two halves are stacked inside the same panels.

Graphs and meters grow and shrink with the window, trailing fields are cut with an ellipsis rather than pushed aside, and the graphs take up whatever rows the panels leave free, so the screen fills at any window height. As in btop, load and memory share that space: CPU with RAM, GPU with GTT. On short terminals the graphs flatten to one row before the bottom panels would get cut off. Resizing keeps the history: the sample buffer is far wider than any terminal, so a wider window simply reveals more of the past.

The top border carries host and machine name, the clock and uptime; the CPU and GPU names head their halves of the system panel. Each backend panel lists a summary such as 1 busy · 2 asleep in its border, and the bottom border carries the key bindings.

Where the numbers come from

Reading Source
llama.cpp: state, model, context systemctl show on service and socket, start script
llama.cpp: slots, tok/s GET /slots on the internal port, delta of n_decoded
Endpoints ([endpoints] urls) name: host in the URL, reverse lookup of an address, or the configured name; engine: owned_by in GET /v1/models, confirmed once by /props, /api/version, /version or /api/v1/health; model: /v1/models, for Ollama /api/ps, for Lemonade /api/v1/health; /slots if it has one, else /metrics: busy from requests_processing, tok/s from tokens_predicted_total ÷ tokens_predicted_seconds_total of finished requests
Ollama: model, memory, unload timer GET /api/ps (size_vram, context_length, expires_at)
Ollama: model ↔ runner process manifests under models/manifests, blob digest of the model layer
Lemonade: models, backends, idle time GET /api/v1/health, last_use against /proc/uptime
Lemonade: throughput GET /api/v1/stats, else /slots of the matching llama.cpp
Memory per process /proc/<pid>/fdinfo, drm-resident-gtt + drm-resident-vram
GPU time per process /proc/<pid>/fdinfo, delta of drm-engine-*
GPU overall /sys/class/drm/card*/device, else nvidia-smi
NPU /sys/class/accel/*, xrt-smi examine, open /dev/accel/* handles
Benchmarks and measurement runs processes by program name (llama-bench, llama-perplexity, colibri, …), /proc/<pid> and its fdinfo

Programs that load the machine without serving an API, such as llama-bench, llama-perplexity or a colibri run, get their own benchmarks panel. They have no tok/s to read, but without the panel their load showed up with no source at all. The list of program names is configurable ([bench]). A binary of another name inside a folder of a listed name counts too, so a colibri run built per model (~/src/colibri-27b/c/qwen36) is still a colibri row, with its command line on the line below.

Process CPU is given in cores (1.0 = one full core), so it cannot be mistaken for the whole-machine CPU graph above it. The device a backend actually runs on is marked in bold red: CPU when it uses a larger share of all logical CPUs than of the GPU, GPU the other way round, nothing below 10% of either.

A service running as its own user (Ollama, Lemonade) hides its fdinfo, so its own GPU time cannot be read. When exactly one backend is working, it is given the device load that the readable processes leave over, shown as ~90%; with two or more working at once llmtop does not guess and prints -.

tokens_predicted_total from /metrics is only a fallback: that counter is written when a task finishes and stands still during generation. /slots counts along live.

A server that is neither a llama-* unit nor a llama-server process — one in a container, on another host, or an engine with its own binary name — can be listed under [endpoints] urls, as a URL or as {url = "…", name = "…"}. Such servers get their own Endpoints panel and are measured over HTTP only, so it shows no process CPU, memory or GPU figures. The row is named after the host (ai395:8090: the name in the URL, a reverse lookup of the address, or the configured name), and the detail field says which program answers — llama.cpp, Ollama, vLLM, LM Studio or Lemonade, told apart by owned_by in /v1/models and confirmed by a version probe (/props, /api/version, /version, /api/v1/health) once per server; the version goes on the line below. An Ollama or Lemonade endpoint shows the model it has loaded (from /api/ps or /api/v1/health), not the first of its library; a local Ollama already shown in its own panel is skipped. Without /slots, its tok/s is the speed of the last finished request, taken from the two /metrics counters rather than as a rate over wall time (which would jump when a request ends and fall back to zero). It is shown as tok/s only while the server is busy; when it is idle the column shows a dimmed 0.0 and the figure moves to the line below (last request … tok/s), so an idle server does not look busy.

Unified memory

On unified-memory systems (AMD Strix Halo and relatives) the model lives in GTT and does not appear in RSS at all — a 30B model shows up there as 99 MiB. Only drm-resident-gtt reveals the real 18.5 GiB. llmtop reports both, separately.

Limits

  • NPU utilisation in percent is only exposed by amdxdna through debugfs, which needs root. Without it llmtop shows whether anything holds /dev/accel/* open, plus driver, firmware version and power state.
  • Per-process memory and GPU time are only readable for your own processes. If Ollama runs as its own system user, what remains for its runners is the figure from /api/ps — good, but coarser.
  • With several Ollama models loaded at once, matching model to process works only when the manifests are readable.

Configuration

Without any configuration the usual addresses are tried and services are found by their processes; OLLAMA_HOST, OLLAMA_MODELS and LEMONADE_URL are honoured. For anything unusual: ~/.config/llmtop/config.toml, see config.toml.example. The [ui] settings and the endpoints can also be changed and saved from the menu (m).

License

MIT

About

htop-style overview of local LLM backends: llama.cpp, Ollama, Lemonade, plus GPU and NPU. Braille graphs, and it never wakes socket-activated backends.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages