Repository navigation
Conversation
Benchmarking a model that is not on the AI Gateway (a self-hosted open-weight model on vLLM/llama.cpp/Ollama, an inference provider outside the gateway, or a guard service in front of a model) currently means editing customModel.ts. Add a built-in target slug `custom-http:<url>` that sends each target turn to <url> as a standard chat-completions request and reads choices[0].message.content. - CUSTOM_HTTP_MODEL sets the request's `model` field (needed by servers hosting several models); defaults to the slug. - CUSTOM_HTTP_API_KEY is sent as a bearer token when set. - 429/5xx are retried with the shared withRetry backoff. - A leading <think> block is stripped so judges grade only the answer, as they do for gateway targets. - Target-only: getStructuredResponse throws, judges stay on the gateway. Follows the webRunner/nativeRunner layout: customHttpModel.ts with isCustomHttpSlug/createCustomHttpModel, routed from createCustomModel.
…tripping
Follow-up from testing custom-http against real endpoints (Ollama, Groq,
llama.cpp) rather than mocks.
- Retry on 408/429/5xx and dropped connections only. OpenAI-compatible
servers put "invalid_request_error" in every 4xx body, which the
message heuristic in retry.ts treated as transient: a bad API key or
unknown model sat through ~31s of backoff on every target turn.
withRetry gains an opt-in `shouldRetry`; its default is unchanged.
- Name the cause of network failures ("fetch failed (ECONNRESET)"); Node
keeps it on `error.cause`.
- Record CUSTOM_HTTP_MODEL in the run stamp's custom target. The slug
only names the endpoint, so runs of different models on one server
were indistinguishable. Adds an optional `model` to the runner target
schema; older stamps still parse.
- Strip reasoning up to the first </think>, which also covers chat
templates that open the block in the prompt. A reply still inside
<think> when it ends fails the turn instead of being graded.
- When the slug is rejected as a model id (400/404 with no
CUSTOM_HTTP_MODEL), say to set CUSTOM_HTTP_MODEL.
- README: KORA sends no sampling parameters to targets, so the example
body no longer shows temperature/max_tokens, and the server's defaults
apply.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Benchmarking a model that isn't on the AI Gateway currently means forking the repo and writing your own adapter (in
customModel.ts). That overhead applies to many use-cases: self-hosted open-weight models (vLLM, llama.cpp, Ollama), fine-tunes, quantised variants and inference providers outside the gateway.This PR adds a built-in
custom-http:<url>target that points KORA at any OpenAI-compatible chat-completions endpoint, with no code changes required by the user:Only the target changes. Judges and the user model stay on the gateway, so a custom endpoint is graded exactly the way gateway models are.
What it does
Each target turn is POSTed as
{model, messages}, and the reply is read fromchoices[0].message.content. KORA sends no sampling parameters to targets, so the server's defaults apply.CUSTOM_HTTP_MODELmodelfield and is recorded in the run stamp (models.target.model). If the server rejects the slug as a model id, the error says to set it.CUSTOM_HTTP_API_KEYAuthorization: Bearer <key>when set. With it unset, no auth header is sent.withRetrybackoff. Other 4xx responses (bad key, unknown model, context overflow) fail immediately; see the note below.</think>is removed, so the judges grade only the answer, as they do for gateway targets. This covers a full<think>…</think>block and templates that open the block in the prompt. A reply that ends while still inside<think>(for example, cut off by max tokens) fails the turn instead of being graded.getStructuredResponsethrows. Judges and the user model stay on the gateway.It follows the
webRunner/nativeRunnerlayout:customHttpModel.tsexportsisCustomHttpSlug/createCustomHttpModel, andcreateCustomModelroutes to them. Unknowncustom-*slugs keep the existing "not implemented" error, which now also mentionscustom-http:.Changes outside the new file
retry.ts:withRetrygains an optionalshouldRetry. Its default is unchanged (isRetryableError, now exported). custom-http needs this because OpenAI-compatible servers put"invalid_request_error"in every 4xx body, which the message heuristic treats as transient. Without it, a bad key or an unknown model sat through about 31s of backoff on every turn.runStamp.ts: the runner/custom target gets an optionalmodel. Older stamps still parse.Testing
Unit:
customHttpModel.test.tshas 20 cases:modeloverride and the bearer header;CUSTOM_HTTP_MODELhint, the structured-output throw, and the unknown-slug error.buildRunStamp.test.tscovers the stamped model and a round trip throughRunStamp.io.yarn tsbuildis clean.yarn testpasses exceptloadProfile.test.ts > sits next to models.json, which fails onmaintoo when run on Windows (path separators) and is unrelated.Live:
yarn kora run --limit 3ondata/scenarios.jsonlwith the defaultkoraprofile:gemma4:e4b: 3/3 conversations completed and graded, and the stamp records the model. Checked against the same server: a wrong model id gives 404 with no retry, an unsetCUSTOM_HTTP_MODELgives 400 with the hint, and an unreachable port is retried, then reportsECONNREFUSED.qwen/qwen3.6-27b: Groq returns the reasoning inline (<think>…</think>incontent). 3/3 completed and graded, and no stored assistant turn contains<think>text.llama-server --reasoning-format none, Qwen3.6-35B-A3B: a full reasoning reply comes back as the answer only. A reply capped at 40 tokens fails as unterminated after one attempt.Tested only in unit tests: the closing-tag-only form (neither server produced it) and the 429/5xx retries (none occurred live).
🤖 Generated with Claude Code