Skip to content

docs(deploy): self-hosted evaluation guide + accurate GPU requirements - #40

Merged
ghchinoy merged 1 commit into
mainfrom
docs/customer-evaluation-guide
Sep 30, 2026
Merged

ghchinoy merged 1 commit into
mainfrom
docs/customer-evaluation-guide

Conversation

@ghchinoy

Copy link
Copy Markdown
Owner

Summary

Answers three recurring customer questions (hosting TL;DR, required GPU, how to evaluate) with one page and corrects an overstated GPU claim.

New: docs/deploy/evaluate.md — "Evaluate dgem on Your Own GPU"

Ordered steps, each with a pass criterion:

  1. Hardware: RTX PRO 6000 (recommended, tested); L4 works (~3× slower; DISABLE_MM=1 on 32 GiB host RAM, 64 GB with vision on); other GPUs untested.
  2. Readiness: docker run + /health until vllm_ready and warmed; how to mount pre-downloaded weights.
  3. First decision: dgem decide --stats -v samples=1; expect ~55–65 ms Server Denoise on RTX PRO 6000.
  4. Baseline: bench-jev and bench-calibration at -w 4; pass ranges 183–192/231 and 43–45/50 (16 runs each, benchmarks/runs/20260927-image-parity).
  5. Concurrency (optional): same suites at production concurrency, 0 errors.
  6. Custom evaluation: policy → labelled dev/test split (100+ per question) → local dgem serve -u <GPU_HOST> for Studio Batch Eval or the Python runner → scripts/policy_calibration.py → what to record.

Mirrored to docs-site/, added to the "Build, Deploy & Operate" sidebar, linked from docs/index.md, index.mdx, deploy/index.md, deploy/laptop.md and deploy/public-images.md.

Corrections

  • deploy/public-images.md: section heading claimed "Any GPU Host (NVIDIA Blackwell, Ada, Hopper, Ampere)"; we have no receipts for the 4-bit image on Hopper/Ampere. Replaced with a tested/works/untested table. Clarified DISABLE_MM.
  • reference/vertex-vs-cloud-run.md: A100/H100 listed as Vertex machine types now notes the public images are tested only on RTX PRO 6000 and L4.

Verified

  • dgem serve -u http://<non-loopback-host>:8080/v1 alone routes to that host (available_backends: ["cloudrun"]); a loopback URL routes as local. Both work for /api/decide.
  • Gateway does not forward -k tokens to loopback URLs (cmd/serve.go), hence the guide's API_KEY note. Possibly worth a follow-up fix.
  • make docs-sync-check (44/44), scripts/docs_link_check.py (0 broken), make check-public pass; docs-site astro build succeeds (45 pages) and in-page anchors used by the guide resolve.

Not verified

  • The guide's steps were not executed end to end on a fresh self-hosted GPU in this change; the numbers come from existing receipts.

…urately

New page docs/deploy/evaluate.md (mirrored to docs-site, added to the sidebar):
hardware requirements, readiness check, first decision with expected GPU time,
JevBench/calibration baseline with pass ranges from the image parity run
(183-192/231, 43-45/50 over 16 runs), optional concurrency check, and a custom
evaluation on held-out labelled data via a local gateway and
scripts/policy_calibration.py.

public-images.md no longer claims support for Ada/Hopper/Ampere: RTX PRO 6000
is tested, L4 works (~3x slower, DISABLE_MM=1 on 32 GiB), other GPUs untested.
@ghchinoy
ghchinoy merged commit 85e7ae5 into main Sep 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant