docs(deploy): self-hosted evaluation guide + accurate GPU requirements - #40
Merged
Merged
Conversation
…urately New page docs/deploy/evaluate.md (mirrored to docs-site, added to the sidebar): hardware requirements, readiness check, first decision with expected GPU time, JevBench/calibration baseline with pass ranges from the image parity run (183-192/231, 43-45/50 over 16 runs), optional concurrency check, and a custom evaluation on held-out labelled data via a local gateway and scripts/policy_calibration.py. public-images.md no longer claims support for Ada/Hopper/Ampere: RTX PRO 6000 is tested, L4 works (~3x slower, DISABLE_MM=1 on 32 GiB), other GPUs untested.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Answers three recurring customer questions (hosting TL;DR, required GPU, how to evaluate) with one page and corrects an overstated GPU claim.
New:
docs/deploy/evaluate.md— "Evaluate dgem on Your Own GPU"Ordered steps, each with a pass criterion:
DISABLE_MM=1on 32 GiB host RAM, 64 GB with vision on); other GPUs untested.docker run+/healthuntilvllm_readyandwarmed; how to mount pre-downloaded weights.dgem decide --stats -v samples=1; expect ~55–65 ms Server Denoise on RTX PRO 6000.bench-jevandbench-calibrationat-w 4; pass ranges 183–192/231 and 43–45/50 (16 runs each,benchmarks/runs/20260927-image-parity).dgem serve -u <GPU_HOST>for Studio Batch Eval or the Python runner →scripts/policy_calibration.py→ what to record.Mirrored to
docs-site/, added to the "Build, Deploy & Operate" sidebar, linked fromdocs/index.md,index.mdx,deploy/index.md,deploy/laptop.mdanddeploy/public-images.md.Corrections
deploy/public-images.md: section heading claimed "Any GPU Host (NVIDIA Blackwell, Ada, Hopper, Ampere)"; we have no receipts for the 4-bit image on Hopper/Ampere. Replaced with a tested/works/untested table. ClarifiedDISABLE_MM.reference/vertex-vs-cloud-run.md: A100/H100 listed as Vertex machine types now notes the public images are tested only on RTX PRO 6000 and L4.Verified
dgem serve -u http://<non-loopback-host>:8080/v1alone routes to that host (available_backends: ["cloudrun"]); a loopback URL routes aslocal. Both work for/api/decide.-ktokens to loopback URLs (cmd/serve.go), hence the guide'sAPI_KEYnote. Possibly worth a follow-up fix.make docs-sync-check(44/44),scripts/docs_link_check.py(0 broken),make check-publicpass; docs-siteastro buildsucceeds (45 pages) and in-page anchors used by the guide resolve.Not verified