Skip to content

Report instruction and bus-access counts alongside emulated cycles - #107

Merged
sidick merged 1 commit into
mainfrom
clock-mhz-instr-mem-counts
Sep 29, 2026
Merged

sidick merged 1 commit into
mainfrom
clock-mhz-instr-mem-counts

Conversation

@sidick

@sidick sidick commented Sep 29, 2026

Copy link
Copy Markdown
Owner

Follow-up to #103/#104, and the cheap half of what #105 asks for.

--clock-mhz reports emulated time, but time alone doesn't say whether a
workload's time went on arithmetic or on bus traffic — and that distinction
decides whether volamos's timing means anything for that workload. volamos
bills no bus wait states, so its emulated time is close to real hardware for
arithmetic-bound code and very optimistic for bus-bound code: @codewiz measured
1.02x to 25x against cycle-paced Copperline in #105, monotonic in memory
intensity.

This reports the work the run actually did, so a workload can be placed on that
scale from the volamos run by itself, without a second runtime to compare
against:

$ volamos --clock-mhz 25 ./bench
...program output...
volamos: 422158 emulated cycles, 0.016886 s at 25 MHz
volamos: 36581 instructions, 83754 bus accesses (66532 read / 17222 write), 11.54 cycles/instr, 2.29 accesses/instr

This is not wait-state modelling — see the note at the end on why I don't
think that's worth building.

Where the numbers come from

  • Instructions come from the same CycleBatchResult as the cycles, so both
    cover exactly the same span of execution and their ratio is meaningful.
  • Accesses are counted in the m68k::AddressBus impl, not in
    AddressSpace. That separation is what makes the figure mean anything:
    volamos's own native library handlers reach guest memory through
    AddressSpace directly, so their traffic is excluded and what remains is
    what the emulated CPU put on the bus. There's a test pinning exactly that.

Reported only alongside a cycle count, never alone: on the default execution
path the trace JIT reads and writes through the raw fast_mem pointer,
bypassing the AddressBus methods, so the counts would see an arbitrary
fraction of the real traffic — the same reason no cycle count exists there.

Two caveats I found by testing, both documented

I built two deliberately-opposite -O2 loops to check the ratio actually
discriminates. It does, but with a wrinkle worth knowing:

workload accesses/instr writes
byte-copy loop 2.67 819,308
register-only arithmetic loop 1.27 110
  • read includes instruction fetch. The bus methods can't tell a fetch
    from a data read, so accesses/instr has a floor a little above 1.0 rather
    than 0, and it's the margin above that floor that indicates data traffic.
  • write is the clean signal — no fetch component at all, and a 7400x
    spread across those two loops versus 2.1x for the combined ratio. That's why
    reads and writes are reported separately rather than only as a total.

My first attempt at this validation was wrong in a way worth recording: I used
memcpy for the memory-heavy case, which routes to exec's CopyMem — a native
Rust handler, hence free — so the traffic vanished entirely and the two
workloads looked nearly identical. That's the "native handlers are free" caveat
biting in practice, and it's a good argument for this line existing: a
suspiciously low access count is now visible evidence that a workload's real
work is happening outside emulated code.

Changes

  • memory.rs: read/write counters on FlatMemory, plus
    AddressSpace::bus_access_counts with a (0, 0) default so the trait stays
    non-breaking.
  • backend.rs: counters incremented in the AddressBus impl; instructions
    accumulated in run_via_cycles.
  • cpu.rs: Cpu::emulated_instructions, defaulting to 0 like
    emulated_cycles, with the same step-doesn't-accumulate caveat.
  • dispatch.rs: Runtime::emulated_instruction_and_access_counts.
  • main.rs: the second report line, formatting split into a pure
    format_emulated_work so it's testable without capturing stderr.
  • userdocs/CLI-Reference.md: the line, both caveats, and the measured
    discrimination figures.

Testing

cargo test --workspace: 1041 passed, 0 failed (1038 before, + the 3 added
here). Build, clippy --all-targets, and fmt --check clean.

New tests: bus accesses track the CPU and not volamos's own AddressSpace
access; the work line's counts and both ratios, including that reads and writes
stay separate; and no line at all when nothing retired, rather than
divide-by-zero ratios.

Verified end-to-end that both lines are absent without --clock-mhz, present
with it, byte-identical across repeated runs, and that stdout stays clean.

On actually modelling memory speed (#105)

I looked at what that would take and don't recommend it. Three separate gaps:
the m68k crate has no wait-state hook at all (AddressBus::sync is outbound
notification only, and on the 68040 it isn't even called, since
internal_cycles gates on prefetch_enabled()); volamos has no Chip/Fast
region model and ignores MEMF_CHIP/MEMF_FAST for placement; and nothing
models an 040 data cache, which is what block-copy performance on that part
actually is. A flat per-access charge would bill every access at bus rate
where real hardware serves most from cache — replacing an obvious, documented
gap with a plausible-looking number that's wrong by an unknown factor. For the
A/B deltas this is actually used for, systematic bias largely cancels anyway.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EHvk3UiKvgnG2xb71UJ67R

--clock-mhz reports emulated time, but time alone doesn't say whether a
workload's time is arithmetic or bus traffic -- and that distinction
decides whether volamos's timing means anything for it. volamos bills no
bus wait states, so its emulated time is close to real hardware for
arithmetic-bound code and very optimistic for bus-bound code: measured
between 1.02x and 25x against cycle-paced hardware, monotonic in memory
intensity (issue #105).

Adds a second report line carrying instructions retired, guest-CPU bus
reads and writes, and the two derived ratios, so a workload can be placed
on that scale from the volamos run by itself:

  volamos: 422158 emulated cycles, 0.016886 s at 25 MHz
  volamos: 36581 instructions, 83754 bus accesses (66532 read / 17222
           write), 11.54 cycles/instr, 2.29 accesses/instr

Instructions come from the same CycleBatchResult as the cycles, so both
cover the same span of execution. Accesses are counted in the
m68k::AddressBus impl rather than in AddressSpace, which is what makes the
figure meaningful: volamos's own native handlers reach guest memory
through AddressSpace, so their traffic is excluded and what remains is
what the emulated CPU put on the bus.

Reported only alongside a cycle count. On the default path the JIT's
fast_mem pointer bypasses the AddressBus methods entirely, so the counts
would see an arbitrary fraction of real traffic.

Measured while testing, and documented: `read` includes instruction fetch,
which the bus methods cannot distinguish from a data read, so
accesses/instr has a floor above 1.0 and it is the margin above it that
indicates data traffic (byte-copy loop 2.67, register-only arithmetic loop
1.27). `write` carries no fetch component and is the clean signal (819308
vs 110 across those same two loops) -- which is why reads and writes are
reported separately rather than only as a total.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EHvk3UiKvgnG2xb71UJ67R
@sidick
sidick merged commit 584dc59 into main Sep 29, 2026
9 checks passed
@sidick
sidick deleted the clock-mhz-instr-mem-counts branch September 29, 2026 20:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant