Repository navigation
R7: implement CPU memory-wall study harness - #8
Conversation
There was a problem hiding this comment.
Sorry @EmergentMonk, you've used your own review budget of 250,000 diff characters for the last 7 days.
You can request another review in 6 days and 16 hours by commenting @sourcery-ai review. Upgrade to get a review now.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
@codex Please review exact head 10e9a30 against the R7 roadmap and existing RUNE contracts. Report only actionable correctness/evidence defects with minimal reproductions. Pay particular attention to comparison parity, lifecycle timing completeness, modelled byte accounting, observation schema, timer semantics, local evidence-bundle provenance, CI-vs-performance-evidence separation, and any accidental runtime ABI change or unsupported causal claim. Distinguish executed from statically inferred findings. |
Reviewer's GuideThis PR advances the project to R7 by adding a local-first C99 harness for eight CPU memory-wall comparisons, with exact-result parity checks, lifecycle-aware clock() observations, structured TSV output, and reproducible local evidence bundles; CI only validates build, correctness parity, and schema shape, while project documentation and policy explicitly prohibit treating timings as performance or causal hardware evidence. Sequence diagram for an R7 lifecycle-aware observationsequenceDiagram
participant Harness as rune_r7_study
participant Timer as C99 clock()
participant Workload as Comparison variant
participant Output as TSV output
Harness->>Timer: clock()
Harness->>Workload: setup allocation and initialization
Harness->>Timer: clock()
Harness->>Workload: execute comparison workload
Workload-->>Harness: result_u64
Harness->>Timer: clock()
Harness->>Workload: teardown and release
Harness->>Timer: clock()
Harness->>Output: r7_emit(observation)
Note over Harness,Output: Pairwise variants are emitted only after exact result identity matches
Flow diagram for R7 comparison validation and captureflowchart TD
Start["Select working-set size and repeat"] --> Compare["Run comparison variants"]
Compare --> Identity{"Exact result identity matches?"}
Identity -- No --> Fail["Fail closed"]
Identity -- Yes --> Record["Record lifecycle, state, byte, result, and clock fields"]
Record --> Emit["Emit numeric TSV row"]
Emit --> More{"More comparisons or sizes?"}
More -- Yes --> Compare
More -- No --> Done["Raw observation capture complete"]
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
|
@codex Please review exact current head 27acc7c against the R7 study contract. Report only novel actionable correctness/evidence defects. Pay special attention to optimizer-resistant variant separation, clock wrap/error handling, lifecycle timing completeness, exact parity gates, environment-bundle provenance, and the prohibition on treating CI timings as performance evidence. |
|
@codex Please review exact current head 9091c1c against the R7 study contract. Report only novel actionable correctness/evidence defects; pay particular attention to optimizer-resistant variant separation, lifecycle timing, observation accounting, local evidence provenance, and the explicit CI-is-not-performance-evidence boundary. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9091c1cfd2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex Please review exact current head c7d641f against the R7 study/evidence contract. Report only novel actionable correctness or provenance defects; do not repeat the resolved selected-tile lifecycle, untracked dirty-state, CPPFLAGS provenance, immutable destination, evidence-class, compiler identity, or cross-platform memory-profile findings without a new failing case. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c7d641f232
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5aa7173d66
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex Please review exact current head 187c894 against the R7 study/evidence contract. Report only novel actionable correctness or provenance defects; do not repeat the resolved AR toolchain binding, shell-safe command serialization, or local-profile five-repeat-default findings without a new failing case. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 187c894cd6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d60e87899e
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d0db06a611
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 58a8d8a4c6
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dab33111b9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex Please Review this exact SHA against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 5e805f98c9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex Please Review this exact SHA against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 380afca9f0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| for flag_token in $flag_value; do | ||
| case "$flag_token" in |
There was a problem hiding this comment.
Reject GCC subprocess wrappers in recorded flags
Executed reproduction: create a wrapper that delegates every GCC subprocess but invokes cc1 with -include /tmp/r7-header.h, redefine CLOCKS_PER_SEC to 454545 in that header, and run CFLAGS='-wrapper /tmp/r7-wrapper' RUNE_R7_REPEATS=1 CC=/usr/bin/gcc AR=/usr/bin/ar scripts/r7-run-local.sh /tmp/r7-bundle. GCC's -wrapper option wraps its subprocesses; the capture exited successfully, recorded only the wrapper path, and every observation reported 454545. Expected the external compiler wrapper to be rejected before Make or identity-bound; actual evidence depends on mutable, unchecksummed wrapper/header contents and can diverge or fail during replay. Add -wrapper and its argument forms to this token validation.
AGENTS.md reference: AGENTS.md:L233-L233
Useful? React with 👍 / 👎.
| if [ -n "${active_study_pid:-}" ]; then | ||
| kill "-$forwarded_signal" "$active_study_pid" >/dev/null 2>&1 || : | ||
| fi |
There was a problem hiding this comment.
Exit when termination arrives before the study starts
Fresh executed case beyond the fixed active-study signal path: start a one-repeat capture, wait only for build/r7-source-* to appear, then send SIGTERM to the wrapper PID. Expected prompt termination and cleanup, but because active_study_pid is still empty this handler merely records status and returns; the process remained alive for about 9.3 seconds, built the executable, ran the entire study, and only then exited 143. With a larger repeat count, cancellation during snapshot/build can therefore launch the full expensive measurement; exit immediately when no study child exists, or check termination_status before starting it.
Useful? React with 👍 / 👎.
| shell_quote "$study_executable" | ||
| printf " --profile local --repeats " | ||
| shell_quote "$repeats" | ||
| printf " > " | ||
| shell_quote "$out_abs/observations.tsv" |
There was a problem hiding this comment.
Preserve the evidence bundle when replay execution fails
Fresh executed case beyond the fixed early replay-failure path: capture a one-repeat bundle, then run (ulimit -v 120000; sh bundle/command.txt). Expected the failed replay to leave the checksum-bound evidence intact and clean its temporary snapshot; actual execution exited 1 during the study after replacing the original 140-row observations.tsv with a 122-row partial file, made sha256sum -c SHA256SUMS fail, and left the read-only source snapshot behind because set -e skipped the cleanup commands. Redirect replay output to a temporary sibling and replace observations.tsv only after success, with trap-based snapshot cleanup on failure.
AGENTS.md reference: AGENTS.md:L169-L169
Useful? React with 👍 / 👎.
| else | ||
| hash_path=$(resolve_provenance_tool shasum 2>/dev/null || :) | ||
| [ -n "$hash_path" ] || fail "no trusted SHA-256 utility available" | ||
| hash_mode=shasum |
There was a problem hiding this comment.
Sanitize Perl injection variables for the shasum fallback
On hosts where sha256sum is absent and this fallback selects the Perl-based shasum, inherited PERL5OPT/PERL5LIB can execute caller-provided code inside the supposedly trusted hashing utility. The full fallback capture is statically inferred; locally, PERL5LIB=/tmp/hook PERL5OPT=-MR7Hook /usr/bin/shasum -a 256 /etc/hosts executed the module's BEGIN block before hashing. Such a module can replace observations.tsv before checksum generation and force -c verification to succeed, producing finalized evidence whose observations were not emitted by the study. Expected the fixed absolute utility to compute hashes without ambient code injection; clear Perl interpreter injection/search variables before identity capture and every fallback invocation.
AGENTS.md reference: AGENTS.md:L215-L215
Useful? React with 👍 / 👎.
|
@codex Please Review this exact SHA against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2bce8d1749
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| -wrapper|--wrapper|-wrapper=*|--wrapper=*) | ||
| IFS=$old_ifs | ||
| fail "$flag_name must not contain GCC subprocess wrapper controls; external compiler wrappers are not bound evidence inputs" |
There was a problem hiding this comment.
Reject quoted compiler-control tokens
Executed against this exact SHA with CFLAGS="'-wrapper' /tmp/r7-qwrap", where the wrapper injected a header redefining CLOCKS_PER_SEC to 464646: the raw token matcher did not recognize the quoted '-wrapper', Make's recipe shell removed the quotes, and capture finalized with all 140 observations reporting 464646. Expected the documented GCC wrapper prohibition to reject this before Make; actual behavior produces checksum-valid evidence dependent on mutable, unrecorded wrapper/header contents. This is a fresh failing case beyond the fixed unquoted -wrapper form; reject shell quoting/escaping around sensitive tokens or parse the effective shell argv before validation.
AGENTS.md reference: AGENTS.md:L228-L230
Useful? React with 👍 / 👎.
| unset GIT_DIR | ||
| unset GIT_WORK_TREE | ||
| unset GIT_INDEX_FILE | ||
| unset GIT_OBJECT_DIRECTORY | ||
| unset GIT_ALTERNATE_OBJECT_DIRECTORIES |
There was a problem hiding this comment.
Isolate Git provenance checks from configured hooks
Executed against this exact SHA by creating a core.fsmonitor hook that replaces the freshly built build/r7-source-*/build/r7-evidence-*/rune_r7_study with a script printing 999, then running capture with GIT_CONFIG_COUNT=1 GIT_CONFIG_KEY_0=core.fsmonitor GIT_CONFIG_VALUE_0=/tmp/r7-fsmon-hook: the post-build git status validation executed the hook, capture exited 0, and the checksum-valid bundle contained only 999 in observations.tsv. Expected the recorded revision and build configuration to determine the measured executable; actual behavior lets unrecorded environment, global, or repository Git configuration execute code during validation and replace it. Clear Git configuration-injection variables and disable hook-bearing configuration such as core.fsmonitor for provenance commands.
AGENTS.md reference: AGENTS.md:L211-L212
Useful? React with 👍 / 👎.
| shell_quote "$bundle_mkdir_path" | ||
| printf " " | ||
| shell_quote "$source_snapshot" | ||
| printf '\n' |
There was a problem hiding this comment.
Use a replay-specific source snapshot
Executed against this exact SHA by capturing a one-repeat bundle and launching two sh bundle/command.txt processes concurrently: expected the independent replays to complete without sharing mutable build state, but both exited 1—one failed because the hard-coded snapshot directory already existed, and its EXIT cleanup deleted that directory while the other replay was extracting into it, causing the other to fail with missing Makefile, src, include, and study paths. The observations remained checksum-valid, but neither replay succeeded. Qualify the replay snapshot and its build directory with the replay process PID, as is already done for the temporary observations file.
Useful? React with 👍 / 👎.
|
@codex Please Review this exact SHA against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 576d23e396
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| [ -z "$ignored_inputs" ] || | ||
| fail "ignored untracked files exist under compiler-input trees; remove them before evidence capture: $ignored_inputs" | ||
|
|
||
| git_status=$("$git_path" -C "$repo_root" status --porcelain --untracked-files=all) || |
There was a problem hiding this comment.
Force Git status to inspect the repository worktree
When the repository-local config sets core.worktree to another clean checkout, this command inspects that checkout rather than repo_root. Executed against this exact SHA: populate a clean /tmp/foreign, run git config core.worktree /tmp/foreign, create an untracked file in the actual repository, and invoke a one-repeat capture; expected rejection for the dirty repository, but the script finalized a checksum-valid bundle recording working_tree_dirty=false. Force the provenance commands to use repo_root as the worktree or reject a mismatched configured worktree.
AGENTS.md reference: AGENTS.md:L184-L185
Useful? React with 👍 / 👎.
| flag_value=$2 | ||
|
|
||
| case "$flag_value" in | ||
| *'$'*|*'`'*|*';'*|*'&'*|*'|'*|*'<'*|*'>'*|*'#'*|*'~'*|*'*'*|*'?'*|*'['*|*']'*|*'@'*) |
There was a problem hiding this comment.
Reject brace-expanded compiler controls
On hosts where Make's /bin/sh is Bash, braces that pass this filter are expanded before compiler execution and can synthesize forbidden controls. The affected-host capture is statically inferred; I executed /bin/bash --posix with CFLAGS='-{specs=/tmp/r7.specs,specs=/tmp/r7.specs}', which expanded to two -specs=/tmp/r7.specs arguments, and GCC compiled a probe with an injected CLOCKS_PER_SEC=474747; the wrapper itself reached Make rather than rejecting the value. Expected rejection before Make, but a Bash-backed capture can load the external specs file and finalize observations dependent on it. Local gcc --help describes -specs=<file> as overriding built-in specs with that file's contents, so reject brace-expansion syntax alongside the other shell metacharacters.
AGENTS.md reference: AGENTS.md:L230-L230
Useful? React with 👍 / 👎.
| PATH="$build_path" MAKEFLAGS= GNUMAKEFLAGS= MFLAGS= MAKEFILES= MAKEOVERRIDES= \ | ||
| "$make_path" -C "$source_snapshot" -f Makefile "$study_target" \ | ||
| BUILD_DIR="$build_dir_rel" \ | ||
| CC="$cc_path" \ |
There was a problem hiding this comment.
Disable Clang's implicit default configuration
Clang automatically loads a default clang.cfg beside its installed executable even when no --config flag is present, so this build can still consume unbound configuration after the explicit config controls were rejected. Executed against this exact SHA: place -include /tmp/r7-default-clang-header.h in /usr/lib/llvm-20/bin/clang.cfg, redefine CLOCKS_PER_SEC to 484848 in that header, and capture with CC=/usr/bin/clang-20; the checksum-valid bundle reported 484848, while deleting the config/header and running command.txt reported 1000000. Expected the recorded build configuration to reproduce the same compiler inputs; actual metadata only mentioned the config path in version output and did not bind its contents. Local clang --help-hidden documents --no-default-config as disabling default configuration loading, so disable implicit configs or reject and bind any detected one. This is fresh evidence beyond the fixed explicit --config case.
AGENTS.md reference: AGENTS.md:L225-L229
Useful? React with 👍 / 👎.
|
@codex Please Review this exact SHA fa111b1 against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking. |
|
Codex Review: Didn't find any major issues. Another round soon, please! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
@codex Please Review this exact SHA fa111b1 against the existing contract. Report only actionable correctness defects, with a minimal reproduction, expected versus actual behavior, and affected lines. State whether each reproduction was executed or statically inferred. Don’t repeat fixed findings without a new failing case. Keep architectural suggestions separate and non-blocking. |
|
Codex Review: Didn't find any major issues. What shall we delve into next? Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Authority transition clarification
This PR is to be reviewed as the proposed R6 → R7 authority transition.
The R6 values on the base branch are the pre-transition authority under review; they are intentionally superseded by the PR-head R7 values if the transition is accepted. Reviewers should therefore validate the migration itself and the preservation of frozen R1–R6 semantics, rather than enforce the base R6 phase fields against the proposed R7 head.
At the PR head:
CURRENT_PHASE=R7RUNTIME_IMPLEMENTATION_ALLOWED=falseUntil merge,
mainremains R6 authority.Summary
Implements the RUNE roadmap R7 — CPU memory-wall study harness on top of merged R6.
Study comparisons
The local-first harness covers:
Observation contract
Each numeric TSV row records:
Pairwise comparisons fail closed unless exact result identity matches.
Timing semantics
Performance gate
This PR permits raw observation capture only.
Local evidence bundle
Adds scripts/r7-run-local.sh, which captures:
CI
Adds make r7-study-smoke:
Scope
R7 is study tooling only:
Summary by Sourcery
Establish the R7 CPU memory-wall study harness and fail-closed local evidence workflow without changing the runtime ABI or enabling performance claims.
New Features:
Bug Fixes:
Enhancements:
Build:
CI:
Documentation:
Tests:
Chores: