Skip to content

[feat] custom-http:<url> target for OpenAI-compatible endpoints - #40

Open
JRM-IT wants to merge 2 commits into
korabench:mainfrom
JRM-IT:feat/custom-http-model
Open

JRM-IT wants to merge 2 commits into
korabench:mainfrom
JRM-IT:feat/custom-http-model

Conversation

@JRM-IT

@JRM-IT JRM-IT commented Sep 30, 2026 •

Copy link
Copy Markdown

Benchmarking a model that isn't on the AI Gateway currently means forking the repo and writing your own adapter (in customModel.ts). That overhead applies to many use-cases: self-hosted open-weight models (vLLM, llama.cpp, Ollama), fine-tunes, quantised variants and inference providers outside the gateway.

This PR adds a built-in custom-http:<url> target that points KORA at any OpenAI-compatible chat-completions endpoint, with no code changes required by the user:

# Open-weight model on a local vLLM server
CUSTOM_HTTP_MODEL=Qwen/Qwen3-8B \
  yarn kora run custom-http:http://localhost:8000/v1/chat/completions

# Hosted provider that needs an API key
CUSTOM_HTTP_MODEL=openai/gpt-oss-20b CUSTOM_HTTP_API_KEY=$GROQ_API_KEY \
  yarn kora run custom-http:https://api.groq.com/openai/v1/chat/completions

Only the target changes. Judges and the user model stay on the gateway, so a custom endpoint is graded exactly the way gateway models are.

What it does

Each target turn is POSTed as {model, messages}, and the reply is read from choices[0].message.content. KORA sends no sampling parameters to targets, so the server's defaults apply.

CUSTOM_HTTP_MODEL Sets the request's model field and is recorded in the run stamp (models.target.model). If the server rejects the slug as a model id, the error says to set it.
CUSTOM_HTTP_API_KEY Sent as Authorization: Bearer <key> when set. With it unset, no auth header is sent.
Retries 408, 429, 5xx and dropped connections go through the shared withRetry backoff. Other 4xx responses (bad key, unknown model, context overflow) fail immediately; see the note below.
Reasoning Everything up to the first </think> is removed, so the judges grade only the answer, as they do for gateway targets. This covers a full <think>…</think> block and templates that open the block in the prompt. A reply that ends while still inside <think> (for example, cut off by max tokens) fails the turn instead of being graded.
Target-only getStructuredResponse throws. Judges and the user model stay on the gateway.

It follows the webRunner/nativeRunner layout: customHttpModel.ts exports isCustomHttpSlug / createCustomHttpModel, and createCustomModel routes to them. Unknown custom-* slugs keep the existing "not implemented" error, which now also mentions custom-http:.

Changes outside the new file

  • retry.ts: withRetry gains an optional shouldRetry. Its default is unchanged (isRetryableError, now exported). custom-http needs this because OpenAI-compatible servers put "invalid_request_error" in every 4xx body, which the message heuristic treats as transient. Without it, a bad key or an unknown model sat through about 31s of backoff on every turn.
  • runStamp.ts: the runner/custom target gets an optional model. Older stamps still parse.

Testing

Unit: customHttpModel.test.ts has 20 cases:

  • routing, including the model override and the bearer header;
  • missing URL, request body shape, and omission of unset sampling parameters;
  • reasoning stripping (full block, closing tag only, plain passthrough) and the unterminated-block failure;
  • retry on 408/429/500/503, no retry on 400/401/404, and network retry with the cause in the message;
  • missing content, the CUSTOM_HTTP_MODEL hint, the structured-output throw, and the unknown-slug error.

buildRunStamp.test.ts covers the stamped model and a round trip through RunStamp.io. yarn tsbuild is clean. yarn test passes except loadProfile.test.ts > sits next to models.json, which fails on main too when run on Windows (path separators) and is unrelated.

Live: yarn kora run --limit 3 on data/scenarios.jsonl with the default kora profile:

  • Ollama, gemma4:e4b: 3/3 conversations completed and graded, and the stamp records the model. Checked against the same server: a wrong model id gives 404 with no retry, an unset CUSTOM_HTTP_MODEL gives 400 with the hint, and an unreachable port is retried, then reports ECONNREFUSED.
  • Groq, qwen/qwen3.6-27b: Groq returns the reasoning inline (<think>…</think> in content). 3/3 completed and graded, and no stored assistant turn contains <think> text.
  • llama.cpp llama-server --reasoning-format none, Qwen3.6-35B-A3B: a full reasoning reply comes back as the answer only. A reply capped at 40 tokens fails as unterminated after one attempt.

Tested only in unit tests: the closing-tag-only form (neither server produced it) and the 429/5xx retries (none occurred live).

🤖 Generated with Claude Code

JRM-IT and others added 2 commits September 30, 2026 17:24
Benchmarking a model that is not on the AI Gateway (a self-hosted
open-weight model on vLLM/llama.cpp/Ollama, an inference provider outside
the gateway, or a guard service in front of a model) currently means
editing customModel.ts.

Add a built-in target slug `custom-http:<url>` that sends each target turn
to <url> as a standard chat-completions request and reads
choices[0].message.content.

- CUSTOM_HTTP_MODEL sets the request's `model` field (needed by servers
  hosting several models); defaults to the slug.
- CUSTOM_HTTP_API_KEY is sent as a bearer token when set.
- 429/5xx are retried with the shared withRetry backoff.
- A leading <think> block is stripped so judges grade only the answer, as
  they do for gateway targets.
- Target-only: getStructuredResponse throws, judges stay on the gateway.

Follows the webRunner/nativeRunner layout: customHttpModel.ts with
isCustomHttpSlug/createCustomHttpModel, routed from createCustomModel.
…tripping

Follow-up from testing custom-http against real endpoints (Ollama, Groq,
llama.cpp) rather than mocks.

- Retry on 408/429/5xx and dropped connections only. OpenAI-compatible
  servers put "invalid_request_error" in every 4xx body, which the
  message heuristic in retry.ts treated as transient: a bad API key or
  unknown model sat through ~31s of backoff on every target turn.
  withRetry gains an opt-in `shouldRetry`; its default is unchanged.
- Name the cause of network failures ("fetch failed (ECONNRESET)"); Node
  keeps it on `error.cause`.
- Record CUSTOM_HTTP_MODEL in the run stamp's custom target. The slug
  only names the endpoint, so runs of different models on one server
  were indistinguishable. Adds an optional `model` to the runner target
  schema; older stamps still parse.
- Strip reasoning up to the first </think>, which also covers chat
  templates that open the block in the prompt. A reply still inside
  <think> when it ends fails the turn instead of being graded.
- When the slug is rejected as a model id (400/404 with no
  CUSTOM_HTTP_MODEL), say to set CUSTOM_HTTP_MODEL.
- README: KORA sends no sampling parameters to targets, so the example
  body no longer shows temperature/max_tokens, and the server's defaults
  apply.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Thibaut-Fatus

Copy link
Copy Markdown
Collaborator

@Clemalfroy

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants