Skip to content

feat(forge): SkillOpt-style epoch training with mini-batch validation - #28

Merged
criptogus merged 1 commit into
mainfrom
claude/pensive-gauss-M17ZI
May 28, 2026
Merged

criptogus merged 1 commit into
mainfrom
claude/pensive-gauss-M17ZI

Conversation

@criptogus

Copy link
Copy Markdown
Owner

Summary

Brings ideas from microsoft/SkillOpt ("train agent skills like neural nets — epochs, mini-batches, validation gates, without touching weights") into our existing SkillForge loop. Our autoLearnPipeline already had evolutionary search (generations + elite archive + A/B fitness), so this PR adds the missing loop-level training mechanics rather than importing the Python project.

Phase 1 — epoch training + mini-batch validation (runForgeLoop)

  • New inputs: epochs (1–6), batch_size, candidates_per_gen.
  • Each epoch runs eval → auto-learn → validates the candidate patch on a rotating mini-batch of golden cases.
  • Maintains a single best_skill snapshot (the running champion); improvements compound across epochs.
  • Promotion is gated: the champion is only persisted as a new version if it does not regress on the full golden set (loop:after).
  • Per-epoch history + training metadata returned and stored in evolution_trace.

Phase 2 (groundwork) — trajectory-driven signal

  • buildExecutionSignal summarizes real agent runs from skill_executions (success rate, top error kinds, per-model success) and feeds them into autoLearnPipeline as the strongest failure evidence for root-cause analysis and patch proposals.
  • Note: full step-by-step trajectory capture (input/output per step) would require a schema migration; this uses the outcome telemetry available today.

Notes

  • Backwards compatible: defaults (epochs=1) reproduce the previous single-pass behaviour. patch in the result is now nullable (no scorable candidate) — UI guarded.

Test plan

  • Run forge loop with epochs=1 on a package — verify before/after parity with prior behaviour.
  • Run with epochs>1 — verify epoch_history, mini-batch validation, and best-snapshot promotion.
  • Verify a regressing champion is rejected by the full-set gate.
  • Confirm execution telemetry appears in evolution_trace when skill_executions exist.

https://claude.ai/code/session_01EntkmBiYh381pKqvvSFBkg


Generated by Claude Code

Adds neural-network-style training controls to the forge loop, inspired by
microsoft/SkillOpt: iterate eval -> learn -> validate over multiple epochs,
each validating a candidate patch on a rotating mini-batch of golden cases,
keeping a single best_skill snapshot promoted only after a full-set gate.

Also feeds real execution telemetry (skill_executions) into auto-learn as a
trajectory-driven failure signal for root-cause analysis and patch proposals.

https://claude.ai/code/session_01EntkmBiYh381pKqvvSFBkg
@criptogus
criptogus marked this pull request as ready for review May 28, 2026 18:22
@criptogus
criptogus merged commit 2431fa1 into main May 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants