Skip to content

[Enhancement]: Rank the model-visible skill catalog per request with the classifier (skills counterpart of #16181) #16563

Description

@marbence101

What features would you like to see added?

Pick the entries of the model-visible skill catalog per request by relevance to the user's message, using the classification provider from #16180, the same way #16181 selects deferred tools.

Today resolveSkillCatalog (packages/api/src/agents/skills.ts) fills the catalog with the first SKILL_CATALOG_LIMIT = 100 active skills in listing order. endpoints.agents.skills.maxCatalogSkills can lower that number but not change which skills are picked, and skills past the cap also leave activeSkills, so the model cannot invoke them even by exact name. With a synced library of ~1,100 skills (K-Dense scientific-agent-skills, academic-research-skills and Auto-Empirical-Research-Skills via skillSync) this means:

  • roughly 9 in 10 skills are never visible to the model, and which 100 are visible is arbitrary, so the relevant skill is usually missing;
  • every turn still carries the 100-entry catalog, and at a 200k window formatSkillCatalog cuts each description to ~61 characters (🔎 fix: Report Measured Skill Catalog Truncation #15590), too short to choose from;
  • a greeting gets the whole catalog too; we saw a skill whose description says "ALWAYS activate this skill" loaded for a plain "hi".

This is a concrete form of suggestions 3 and 4 in #13127 (dynamic preselection / cheap router model), now that a classification interface exists.

Proposed behavior:

  1. When classification is enabled and the visible catalog is larger than the cap, score every visible skill against the latest user message and list the top N.
  2. Skills below a probability threshold are not listed; if nothing clears it, the turn gets no catalog (and no skill tool).
  3. A skill the user names in the message is always listed.
  4. activeSkills stays the full active set: ranking changes only the catalog text, and every active skill stays invocable by exact name.
  5. Catalogs that already fit under the cap skip the classifier (no cost for agents with a few selected skills).
  6. Any classifier failure (timeout, 429, auth, billing) falls back to today's behavior with a logged warning; the turn never blocks.

Possible config, next to the existing key:

endpoints:
  agents:
    skills:
      maxCatalogSkills: 20      # existing: entries in the catalog
      selection:
        enabled: true           # uses the `classification` provider from #16180
        minProbability: 0.6
        batchSize: 300
        timeoutMs: 5000

More details

We run this as a patch on v0.8.8-rc2 with TypeSafe's Jev (/v1/systemone). What we learned:

  • One boolean question per skill ("an agent about to work on this task would do it better if it first loaded skill X: <description, first 300 chars>"), with only the user message as state. Following TypeSafe's guidance, the candidate list stays out of the state, since unrelated text there lowers accuracy.
  • Jev answers each question independently of the others in the request, so the catalog can be split into batches that run in parallel, followed by one global sort. The same skill scored alone, in two 200-question batches and in one 400-question batch differed by at most 0.04, so no tournament rounds or keyword prefilter are needed.
  • Ranking needs a scan of the whole accessible catalog, so the page budget (MAX_CATALOG_PAGES = 10) needs its own limit in that mode (we use 50).

Measured with jev-1.13.0 on 1,048 skills (2026-09-29):

requests per turn 4 (300 questions each, at most 8 in flight)
wall time 0.7–1.2 s
classifier input ~105k tokens/turn ≈ $0.0044 at $0.042/M (output is free)
"staggered DiD" top 4 all econometrics / DiD skills at 0.98
"hierarchical Bayesian model in PyMC" pymc, bayesian-estimation at 0.98
off-topic ("a name for my cat") best 0.51, all others ≤ 0.15, so no catalog that turn

With 20 entries the catalog fits the SDK budget without further cuts (up to ~26 entries keep the full 250 characters at a 200k window, per #15590).

Open questions:

Happy to open a PR on top of #16180 if this direction is welcome.

Which components are impacted by your request?

Endpoints

Pictures

No response

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ✨ enhancementNew feature or request🗺️ Backend Platformcodegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions