Who holds the VRAM, who burns the engines, what has spilled to RAM — per process, at a glance, and headless for agents.
Windows answers "how full is the card" but buries "who filled it" behind two right-clicks and a
hidden column — and the number it shows there (WDDM committed bytes) can exceed physical VRAM,
because WDDM virtualizes GPU memory: on the reference box dwm.exe committed 46 GB against a
16 GB card while holding ~1 GB resident. A multi-tenant card (resident LLM + speech model +
browser + compositor + agent fleet) needs the truthful ranked list. That is this tool.
Single ~0.5 MB exe. C/C++ only, OS APIs only (GPU perf counters + DXGI + D3DKMT), zero dependencies, no installer, no elevation, no process handles.
vramtop pretty console snapshot (colors, bars; auto-plain when piped)
vramtop -w live TUI q quit · s sort (vram→sysmem→util→pid→name) · g group · p pause
vramtop --gui live treemap window G group · T pin on top · F5 refresh
vramtopw.exe the same window with no console at all — double-click it, pin it
vramtop -j one JSON line · vramtop -j --no-util = fastest read (~150 ms)
vramtop -j -w -n 2000 NDJSON stream, one snapshot per 2 s (engine % spans the interval)
vramtop --stamp one-line co-tenancy receipt (see below)
vramtop --mcp MCP stdio server — tools: gpu_snapshot, gpu_stamp
vramtop --spool change-gated "lane<TAB>text" stream (fusor/TOWER tailer format)
vramtop --selftest parser, treemap, gate and live-collection checks
vramtop --about the organ's self-description as JSON: verbs, MCP, health — what `peek env` reads
vramtop --help every flag, the JSON shape, exit codes
The window shows the whole card: every tenant sized by resident VRAM, an unattributed
tile for what the driver holds without a process, and a dark free tile — so 16 GB looks like
16 GB. Each adapter band carries the VRAM bar, engine load, temperature / fan / power / memory
clock, and a history sparkline (VRAM as columns, load as the line).
- vram (per process) = bytes resident in VRAM right now — the kernel's
Local Usagecounter. No process is opened for this, so protected processes (csrss) report too. - sysmem = that process's GPU data currently living in system RAM (
Non Local Usage). When the card is full this is the spill; it is the number that explains why a model that fit yesterday is slow today. Shown per row, per adapter, and in every stamp asspill_mib. - A trailing
*means the resident counter had no row for that process and the WDDM commit (Task Manager's column) is shown instead — it can legitimately exceed physical. JSON always carries both:vram_resident_bytes(nullable) andvram_committed_bytes. - Adapter
VRAM x/yis resident usage from the adapter's own counter. Rows do not sum to it exactly (driver/kernel reservations one way, shared surfaces the other); the treemap shows the shortfall asunattributedrather than hiding it. - Engine % needs two counter reads. One-shot reads sleep
--sample-msbetween them (default 500); loops keep the query open, so each frame's % spans the interval since the previous frame.sample_msin JSON is the measured window;0means no load data in that snapshot. - Temperature / fan / power (% of TDP) / memory clock come from
ADAPTERPERFDATA(WDDM 2.4+), vendor-neutral; drivers that don't report (the iGPU) show nothing rather than zeros.
The number that voids GPU timing measurements is the one nobody records: how much VRAM was free
at load, what had already spilled, and who else was on the card. vramtop --stamp is that
receipt — one line, embeddable verbatim in any log or tape row (--json for an object,
--sample-ms 500 to add load):
gpu_stamp t=2026-09-01T14:41:27 RTX_4070_Ti_SUPER free_mib=8491/16063 budget_mib=15295 spill_mib=479 tenants=python:2585,llama-server:1571,nemo-speech:1179,dwm:1076,chrome:344x2,claude:225x2,+18
vramtop --json --no-util # fastest full read
vramtop --stamp # before/after a model load — the mib_free_at_load receipt
claude mcp add vramtop -- C:/GPUz/vramtop.exe --mcp # tools: gpu_snapshot, gpu_stampJSON shape (single line, stable field names):
adapters[] {index,name,vendor,luid,phys,software, vram_total_bytes, vram_used_bytes, vram_budget_bytes, shared_total_bytes, shared_used_bytes, sysmem_used_bytes, util_pct, engines[{name,pct}], temp_c|null, fan_rpm|null, power_pct|null, mem_clock_mhz|null} ·
processes[] {pid,adapter,name,path, vram_resident_bytes|null, vram_committed_bytes, vram_bytes, sysmem_bytes, shared_bytes, util_pct, top_engine} — vram_bytes is resident when
known, else committed. Exit codes: 0 ok · 1 bad args · 2 collector error (JSON still emitted)
· 3 selftest failed.
Collection lives in gpu_snapshot.{h,cpp} — a self-contained organ (a Collector with one
open counter query and cached handles) liftable into C:\tower (src/adapters/) unchanged.
--spool already emits the change-gated lane<TAB>text deltas a fusor tailer eats (transport
v0, the dsh-hop0 spool pattern): a line only when VRAM, spill or load moves past
--gate-mb / --gate-pct or a tenant appears/vanishes; silence stays silent. Wiring it into
towerd's vitals lane is a later, separate step.
build.bat # VS2022: cl /std:c++20 /O2 /W4 /permissive- /utf-8 /MT → vramtop.exe + vramtopw.exe
Files: gpu_snapshot.h/.cpp (collector) · treemap.h (squarified layout, pure) ·
app_util.h (options, formatting, Unicode-width columns) · vramtop.cpp (console modes,
JSON, MCP, stamp, spool, selftest) · vramtop_gui.cpp (the window) · vramtop.rc +
vramtop.manifest (version info, PerMonitorV2 DPI, UTF-8 code page).
- WSL/VM GPU work is attributed to
vmmem*(labeled "(WSL/VM)"), not to the guest process. - Rows can slightly exceed the adapter total (a surface shared by two processes counts for
both);
unattributedis clamped at zero rather than invented. - Windows only, by construction — the whole point is the WDDM accounting path.
GPU-Z, nvidia-smi and friends are card instruments: they ask the driver, and the card has no notion of which Windows PID owns what. Per-process attribution is OS kernel accounting (WDDM / dxgkrnl), a different code path entirely — and under WDDM, nvidia-smi's per-process numbers are absent or under-reported. Datacenter tooling (DCGM, per-container accounting) solves it but doesn't run on Windows desktops; gamers run one workload and never needed it. A consumer card carrying a resident LLM, a speech model, a browser and an agent fleet is a multi-tenant server workload on a desktop — that's the gap this tool sits in.