Skip to content

--clock-mhz: report native-handler call counts at exit (fixes #109) - #110

Merged
sidick merged 1 commit into
mainfrom
clock-mhz-handler-counts
Oct 2, 2026
Merged

sidick merged 1 commit into
mainfrom
clock-mhz-handler-counts

Conversation

@sidick

@sidick sidick commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Native library handlers run in zero emulated cycles under --clock-mhz — documented, but previously invisible: a benchmark had no way to see how much of its work bypassed the cycle count. AmigaPorts/m68k-amigaos-gcc#89 hit this for real: GCC codegen patches A/B-benchmarked on a real A3000 vs volamos --clock-mhz 25 showed a memcpy benchmark improving 5.3% on hardware but 49.1% under volamos — the signature of copies vanishing into exec's native CopyMem.

  • LibraryTable::dispatch counts every dispatched call per slot (one integer hashmap bump on the hot path; library/handler names are resolved only at report time) and accumulates the CopyMem/CopyMemQuick pair's D0 byte size, read before the handler runs so it's the caller's own value.
  • Runtime::native_call_counts() exposes both, gated on a configured clock rate and scoped to the top-level run, exactly like emulated_cycles().
  • The CLI prints a third report line under --clock-mhz (omitted when the guest made no library calls):
volamos: 183 native library calls ran in zero emulated cycles: exec.library/CopyMem x120, dos.library/Write x41, ...; CopyMem/CopyMemQuick moved 491520 bytes natively

Top 5 handlers named, the rest folded into "and N more"; the copy total is one combined clause so the pair's shared number isn't double-reported.

Non-goal (per the issue): billing approximate cycles for native handlers. This makes the unbilled work visible; a synthetic cost model is a separate discussion.

Tests: core coverage that dispatch counts CopyMem+CopyMemQuick and accumulates both calls' bytes, that the accessor is None without a clock rate; CLI formatting tests (top-N + tail fold, zero-copy clause omission, empty case). 1053 tests passing, clippy/fmt clean, strict docs build passes. Docs: CLI-Reference --clock-mhz section (example output + explanation + warning-bullet counterweight), --help text, Changelog 0.8.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GNNXMtYaZVHe8ihLL2GsrY

Native library handlers run in zero emulated cycles -- documented, but
previously invisible: a benchmark had no way to see how much of its
work bypassed the cycle count. Found in the wild via
AmigaPorts/m68k-amigaos-gcc#89, where GCC codegen patches were A/B
benchmarked on a real A3000 vs volamos --clock-mhz 25 and a memcpy
benchmark improved 5.3% on hardware but 49.1% under volamos -- the
signature of copies vanishing into exec's native CopyMem.

LibraryTable::dispatch now counts every dispatched call per slot (one
integer hashmap bump; names resolved only at report time) and
accumulates the CopyMem/CopyMemQuick pair's D0 byte size, read before
the handler runs. Runtime::native_call_counts exposes both under the
same clock-rate gating and top-level-run-only boundary as
emulated_cycles. The CLI prints a third report line under --clock-mhz:

  volamos: 183 native library calls ran in zero emulated cycles:
  exec.library/CopyMem x120, ...; CopyMem/CopyMemQuick moved 491520
  bytes natively

Omitted when the guest made no library calls. Docs: CLI-Reference
--clock-mhz section, the zero-cycle warning bullet, --help text, and a
Changelog 0.8 entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GNNXMtYaZVHe8ihLL2GsrY
@sidick
sidick merged commit 0a68fab into main Oct 2, 2026
9 checks passed
@sidick
sidick deleted the clock-mhz-handler-counts branch October 2, 2026 05:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant