Latency, diversity, and quality benchmark for autoregressive decoding strategies on RTX 2070: greedy, top-k, top-p, min-p, and beam search across GPT-2 models.
-
Updated
Jul 19, 2026 - Python
Latency, diversity, and quality benchmark for autoregressive decoding strategies on RTX 2070: greedy, top-k, top-p, min-p, and beam search across GPT-2 models.
Boost LLM reliability with dynamic sampling, retry logic, and a lightweight Go‑based proxy.
Add a description, image, and links to the min-p topic page so that developers can more easily learn about it.
To associate your repository with the min-p topic, visit your repo's landing page and select "manage topics."