Repository navigation
Report instruction and bus-access counts alongside emulated cycles - #107
Merged
Merged
Conversation
--clock-mhz reports emulated time, but time alone doesn't say whether a workload's time is arithmetic or bus traffic -- and that distinction decides whether volamos's timing means anything for it. volamos bills no bus wait states, so its emulated time is close to real hardware for arithmetic-bound code and very optimistic for bus-bound code: measured between 1.02x and 25x against cycle-paced hardware, monotonic in memory intensity (issue #105). Adds a second report line carrying instructions retired, guest-CPU bus reads and writes, and the two derived ratios, so a workload can be placed on that scale from the volamos run by itself: volamos: 422158 emulated cycles, 0.016886 s at 25 MHz volamos: 36581 instructions, 83754 bus accesses (66532 read / 17222 write), 11.54 cycles/instr, 2.29 accesses/instr Instructions come from the same CycleBatchResult as the cycles, so both cover the same span of execution. Accesses are counted in the m68k::AddressBus impl rather than in AddressSpace, which is what makes the figure meaningful: volamos's own native handlers reach guest memory through AddressSpace, so their traffic is excluded and what remains is what the emulated CPU put on the bus. Reported only alongside a cycle count. On the default path the JIT's fast_mem pointer bypasses the AddressBus methods entirely, so the counts would see an arbitrary fraction of real traffic. Measured while testing, and documented: `read` includes instruction fetch, which the bus methods cannot distinguish from a data read, so accesses/instr has a floor above 1.0 and it is the margin above it that indicates data traffic (byte-copy loop 2.67, register-only arithmetic loop 1.27). `write` carries no fetch component and is the clean signal (819308 vs 110 across those same two loops) -- which is why reads and writes are reported separately rather than only as a total. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EHvk3UiKvgnG2xb71UJ67R
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to #103/#104, and the cheap half of what #105 asks for.
--clock-mhzreports emulated time, but time alone doesn't say whether aworkload's time went on arithmetic or on bus traffic — and that distinction
decides whether volamos's timing means anything for that workload. volamos
bills no bus wait states, so its emulated time is close to real hardware for
arithmetic-bound code and very optimistic for bus-bound code: @codewiz measured
1.02x to 25x against cycle-paced Copperline in #105, monotonic in memory
intensity.
This reports the work the run actually did, so a workload can be placed on that
scale from the volamos run by itself, without a second runtime to compare
against:
This is not wait-state modelling — see the note at the end on why I don't
think that's worth building.
Where the numbers come from
CycleBatchResultas the cycles, so bothcover exactly the same span of execution and their ratio is meaningful.
m68k::AddressBusimpl, not inAddressSpace. That separation is what makes the figure mean anything:volamos's own native library handlers reach guest memory through
AddressSpacedirectly, so their traffic is excluded and what remains iswhat the emulated CPU put on the bus. There's a test pinning exactly that.
Reported only alongside a cycle count, never alone: on the default execution
path the trace JIT reads and writes through the raw
fast_mempointer,bypassing the
AddressBusmethods, so the counts would see an arbitraryfraction of the real traffic — the same reason no cycle count exists there.
Two caveats I found by testing, both documented
I built two deliberately-opposite
-O2loops to check the ratio actuallydiscriminates. It does, but with a wrinkle worth knowing:
readincludes instruction fetch. The bus methods can't tell a fetchfrom a data read, so accesses/instr has a floor a little above 1.0 rather
than 0, and it's the margin above that floor that indicates data traffic.
writeis the clean signal — no fetch component at all, and a 7400xspread across those two loops versus 2.1x for the combined ratio. That's why
reads and writes are reported separately rather than only as a total.
My first attempt at this validation was wrong in a way worth recording: I used
memcpyfor the memory-heavy case, which routes to exec'sCopyMem— a nativeRust handler, hence free — so the traffic vanished entirely and the two
workloads looked nearly identical. That's the "native handlers are free" caveat
biting in practice, and it's a good argument for this line existing: a
suspiciously low access count is now visible evidence that a workload's real
work is happening outside emulated code.
Changes
memory.rs: read/write counters onFlatMemory, plusAddressSpace::bus_access_countswith a(0, 0)default so the trait staysnon-breaking.
backend.rs: counters incremented in theAddressBusimpl; instructionsaccumulated in
run_via_cycles.cpu.rs:Cpu::emulated_instructions, defaulting to0likeemulated_cycles, with the samestep-doesn't-accumulate caveat.dispatch.rs:Runtime::emulated_instruction_and_access_counts.main.rs: the second report line, formatting split into a pureformat_emulated_workso it's testable without capturing stderr.userdocs/CLI-Reference.md: the line, both caveats, and the measureddiscrimination figures.
Testing
cargo test --workspace: 1041 passed, 0 failed (1038 before, + the 3 addedhere). Build, clippy
--all-targets, andfmt --checkclean.New tests: bus accesses track the CPU and not volamos's own
AddressSpaceaccess; the work line's counts and both ratios, including that reads and writes
stay separate; and no line at all when nothing retired, rather than
divide-by-zero ratios.
Verified end-to-end that both lines are absent without
--clock-mhz, presentwith it, byte-identical across repeated runs, and that stdout stays clean.
On actually modelling memory speed (#105)
I looked at what that would take and don't recommend it. Three separate gaps:
the
m68kcrate has no wait-state hook at all (AddressBus::syncis outboundnotification only, and on the 68040 it isn't even called, since
internal_cyclesgates onprefetch_enabled()); volamos has no Chip/Fastregion model and ignores
MEMF_CHIP/MEMF_FASTfor placement; and nothingmodels an 040 data cache, which is what block-copy performance on that part
actually is. A flat per-access charge would bill every access at bus rate
where real hardware serves most from cache — replacing an obvious, documented
gap with a plausible-looking number that's wrong by an unknown factor. For the
A/B deltas this is actually used for, systematic bias largely cancels anyway.
🤖 Generated with Claude Code
https://claude.ai/code/session_01EHvk3UiKvgnG2xb71UJ67R