Skip to content

Repository files navigation

vramtop — a Task Manager for the GPU

Who holds the VRAM, who burns the engines, what has spilled to RAM — per process, at a glance, and headless for agents.

Windows answers "how full is the card" but buries "who filled it" behind two right-clicks and a hidden column — and the number it shows there (WDDM committed bytes) can exceed physical VRAM, because WDDM virtualizes GPU memory: on the reference box dwm.exe committed 46 GB against a 16 GB card while holding ~1 GB resident. A multi-tenant card (resident LLM + speech model + browser + compositor + agent fleet) needs the truthful ranked list. That is this tool.

Single ~0.5 MB exe. C/C++ only, OS APIs only (GPU perf counters + DXGI + D3DKMT), zero dependencies, no installer, no elevation, no process handles.

vramtop — the card as a treemap: tenants, unattributed, free

Use

vramtop                 pretty console snapshot (colors, bars; auto-plain when piped)
vramtop -w              live TUI   q quit · s sort (vram→sysmem→util→pid→name) · g group · p pause
vramtop --gui           live treemap window   G group · T pin on top · F5 refresh
vramtopw.exe            the same window with no console at all — double-click it, pin it
vramtop -j              one JSON line   ·   vramtop -j --no-util  = fastest read (~150 ms)
vramtop -j -w -n 2000   NDJSON stream, one snapshot per 2 s (engine % spans the interval)
vramtop --stamp         one-line co-tenancy receipt (see below)
vramtop --mcp           MCP stdio server — tools: gpu_snapshot, gpu_stamp
vramtop --spool         change-gated "lane<TAB>text" stream (fusor/TOWER tailer format)
vramtop --selftest      parser, treemap, gate and live-collection checks
vramtop --about        the organ's self-description as JSON: verbs, MCP, health — what `peek env` reads
vramtop --help          every flag, the JSON shape, exit codes

The window shows the whole card: every tenant sized by resident VRAM, an unattributed tile for what the driver holds without a process, and a dark free tile — so 16 GB looks like 16 GB. Each adapter band carries the VRAM bar, engine load, temperature / fan / power / memory clock, and a history sparkline (VRAM as columns, load as the line).

The numbers, honestly

  • vram (per process) = bytes resident in VRAM right now — the kernel's Local Usage counter. No process is opened for this, so protected processes (csrss) report too.
  • sysmem = that process's GPU data currently living in system RAM (Non Local Usage). When the card is full this is the spill; it is the number that explains why a model that fit yesterday is slow today. Shown per row, per adapter, and in every stamp as spill_mib.
  • A trailing * means the resident counter had no row for that process and the WDDM commit (Task Manager's column) is shown instead — it can legitimately exceed physical. JSON always carries both: vram_resident_bytes (nullable) and vram_committed_bytes.
  • Adapter VRAM x/y is resident usage from the adapter's own counter. Rows do not sum to it exactly (driver/kernel reservations one way, shared surfaces the other); the treemap shows the shortfall as unattributed rather than hiding it.
  • Engine % needs two counter reads. One-shot reads sleep --sample-ms between them (default 500); loops keep the query open, so each frame's % spans the interval since the previous frame. sample_ms in JSON is the measured window; 0 means no load data in that snapshot.
  • Temperature / fan / power (% of TDP) / memory clock come from ADAPTERPERFDATA (WDDM 2.4+), vendor-neutral; drivers that don't report (the iGPU) show nothing rather than zeros.

The instrument: --stamp

The number that voids GPU timing measurements is the one nobody records: how much VRAM was free at load, what had already spilled, and who else was on the card. vramtop --stamp is that receipt — one line, embeddable verbatim in any log or tape row (--json for an object, --sample-ms 500 to add load):

gpu_stamp t=2026-09-01T14:41:27 RTX_4070_Ti_SUPER free_mib=8491/16063 budget_mib=15295 spill_mib=479 tenants=python:2585,llama-server:1571,nemo-speech:1179,dwm:1076,chrome:344x2,claude:225x2,+18

Agents

vramtop --json --no-util          # fastest full read
vramtop --stamp                   # before/after a model load — the mib_free_at_load receipt
claude mcp add vramtop -- C:/GPUz/vramtop.exe --mcp     # tools: gpu_snapshot, gpu_stamp

JSON shape (single line, stable field names): adapters[] {index,name,vendor,luid,phys,software, vram_total_bytes, vram_used_bytes, vram_budget_bytes, shared_total_bytes, shared_used_bytes, sysmem_used_bytes, util_pct, engines[{name,pct}], temp_c|null, fan_rpm|null, power_pct|null, mem_clock_mhz|null} · processes[] {pid,adapter,name,path, vram_resident_bytes|null, vram_committed_bytes, vram_bytes, sysmem_bytes, shared_bytes, util_pct, top_engine} — vram_bytes is resident when known, else committed. Exit codes: 0 ok · 1 bad args · 2 collector error (JSON still emitted) · 3 selftest failed.

The TOWER / fusor seam (deferred, shaped)

Collection lives in gpu_snapshot.{h,cpp} — a self-contained organ (a Collector with one open counter query and cached handles) liftable into C:\tower (src/adapters/) unchanged. --spool already emits the change-gated lane<TAB>text deltas a fusor tailer eats (transport v0, the dsh-hop0 spool pattern): a line only when VRAM, spill or load moves past --gate-mb / --gate-pct or a tenant appears/vanishes; silence stays silent. Wiring it into towerd's vitals lane is a later, separate step.

Build

build.bat     # VS2022: cl /std:c++20 /O2 /W4 /permissive- /utf-8 /MT → vramtop.exe + vramtopw.exe

Files: gpu_snapshot.h/.cpp (collector) · treemap.h (squarified layout, pure) · app_util.h (options, formatting, Unicode-width columns) · vramtop.cpp (console modes, JSON, MCP, stamp, spool, selftest) · vramtop_gui.cpp (the window) · vramtop.rc + vramtop.manifest (version info, PerMonitorV2 DPI, UTF-8 code page).

Known limits

  • WSL/VM GPU work is attributed to vmmem* (labeled "(WSL/VM)"), not to the guest process.
  • Rows can slightly exceed the adapter total (a surface shared by two processes counts for both); unattributed is clamped at zero rather than invented.
  • Windows only, by construction — the whole point is the WDDM accounting path.

Why the classics can't do this

GPU-Z, nvidia-smi and friends are card instruments: they ask the driver, and the card has no notion of which Windows PID owns what. Per-process attribution is OS kernel accounting (WDDM / dxgkrnl), a different code path entirely — and under WDDM, nvidia-smi's per-process numbers are absent or under-reported. Datacenter tooling (DCGM, per-container accounting) solves it but doesn't run on Windows desktops; gamers run one workload and never needed it. A consumer card carrying a resident LLM, a speech model, a browser and an agent fleet is a multi-tenant server workload on a desktop — that's the gap this tool sits in.

About

A Task Manager for the GPU on Windows: per-process VRAM (resident vs committed) + engine load. Treemap GUI, live TUI, JSON/NDJSON, MCP server, co-tenancy stamps for agents. C++20, zero deps, single exe.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages