You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
test: prove subprocess and production coverage (Fixes #534) - #562
Make coverage evidence demonstrate real server execution and expose production gaps without changing existing whole-workspace regression gates.
Compare exact idle/info child profiles and require positive execution in the real handler, transport dispatch, and response writer after graceful shutdown.
Add production/test and changed-line diagnostics with conservative accounting for LLVM summary entries without unique source lines.
Require profile proof and reporting in Linux/Windows PR and baseline jobs; add native macOS coverage measured against the exact base on the same runner.
Document classification, branch-instrumentation limits, and artifact semantics.
Validation: 76 Python tests, 19 Windows native cases, workspace formatting/Clippy, and independent review pass. Actual Windows LLVM profiles prove all three execution witnesses increase from 0 to 1. Hosted macOS validation and quality inspection remain pending; no signing/release services run.
Verify exact idle/info child profiles and real handler/transport/writer counter increases. Add conservative production/changed-line diagnostics without altering existing raw coverage gates, and measure native macOS against the exact base on one runner.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
A regression must exceed both the documented absolute and relative budget. Environment and manager inventories must match exactly within the same inventory schema.
A regression must exceed both the documented absolute and relative budget. Environment and manager inventories must match exactly within the same inventory schema.
A regression must exceed both the documented absolute and relative budget. Environment and manager inventories must match exactly within the same inventory schema.
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The new macOS coverage gate still needs hosted validation, and the source scanner has unresolved scaling issues.
Review effort: Balanced Findings: None
Previously missed (2)
In code that hasn't changed since last review
Avoid quadratic source slicing in raw-string detection
scripts/coverage_detail.py:118
code_mask passes source[i:] to this regex on nearly every unmasked character, copying the remaining file each time. That makes production/test classification quadratic across the Rust files processed by summarize. Match at offset i without slicing, and use the match's absolute end to find the raw-string terminator.
Avoid quadratic source slicing in character-literal detection
scripts/coverage_detail.py:137
This character-literal check copies the remaining source for every apostrophe, including Rust lifetime markers. Even after fixing the raw-string scan above, files with many lifetimes or character literals can still take quadratic time to classify. Match at offset i and use the match's absolute end.
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The new cross-platform coverage gate has unresolved reporting issues, and hosted macOS validation remains pending.
Review effort: Balanced Findings: None
Previously missed (2)
In code that hasn't changed since last review
Recognize test-only conditional attributes requiring test
scripts/coverage_detail.py:155
The matcher handles only #[cfg(test)] and #[test], so it treats helpers inside #[cfg(all(test, unix))] modules as production. For example, the executable get_env_var branch at crates/pet-homebrew/src/lib.rs:193 belongs to such a test-only module but contributes to the new production totals on Unix/macOS. Recognize conditional attributes that require test (including all(test, unix)) when excluding an item, and add a classification regression test; avoid treating any(test, unix) as test-only.
Classify unmapped LLVM entries in test files as test coverage
scripts/coverage_detail.py:242
For a source under tests/ or benches/, all mapped lines are classified as tests at line 229, but this line puts every unmapped LLVM LF entry into production_found. Such entries cannot be production in these test-only files, so the production gap is overstated. Put their unmapped entries in test_found instead, still uncovered, and add a regression test using a test file with LF larger than its DA count.
Exclude all/any predicates that require test without excluding optional test branches. Keep unmapped integration-test and benchmark entries in the test denominator while preserving conservative bounds and raw coverage gates.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Current-head update (e79fcac): the source-scanner allocation findings and both classification findings are addressed, with 81 Python tests and independent review. Raw coverage gates remain unchanged.
Native macOS validation has completed successfully: lines 80.561% versus the same-runner exact base's 80.504% (+0.058pp), functions unchanged at 83.180%. The isolated idle/info profiles prove 0 -> 1 execution at all three witnesses: handler, transport dispatch, and writer.
I am also investigating the repeated macOS performance drift rather than accepting only the green budget result. A diagnostic-only eight-pass, counterbalanced same-host comparison of exact main and this head is running here: https://github.com/microsoft/python-environment-tools/actions/runs/36749488202 . It retains full inventories, binary/harness provenance, and all samples. No production Rust or Cargo inputs changed in this PR; the control will help distinguish PR effects from runner variation. This PR remains draft until review and quality inspection are complete. The diagnostic branch is not intended for merge.
The current-head macOS performance failure was investigated rather than ignored: discovery P50 was 205 ms vs 94 ms, exceeding the unchanged absolute/relative gate by 111 ms / 118.1%. Linux and Windows performance improved with matching inventories; their coverage deltas are 0.000 pp / +0.009 pp, and native macOS coverage is +0.058 pp with all isolated subprocess witnesses passing.
The exact-ref same-host control completed successfully: https://github.com/microsoft/python-environment-tools/actions/runs/36749488202. It tested base 597b599 and head e79fcac in B-H-H-B-H-B-B-H order, retained every sample, verified binary hashes before each pass, and all eight inventories have the same SHA-256 (10 environments, 1 manager).
Run-level statistic
Base passes
Head passes
Discovery P50 (ms)
105, 90, 146, 137
127, 112, 124, 104
Discovery P95 (ms)
160, 96, 170, 154
207, 194, 149, 116
Refresh RTT P50 (ms)
106, 91, 147, 138
128, 113, 124, 105
Cold RTT P50 (ms)
274, 208, 358, 323
324, 275, 274, 229
Median run-level discovery P50 is 121 ms base / 118 ms head; refresh P50 122 / 118.5 ms. The tail remains noisy (median discovery P95 157 / 171.5 ms) and is not being discarded. Both original jobs used the same macOS image/architecture; this is not an image-transition claim. The controlled distributions support run/host variance, not a stable source regression.
I am making one targeted confirmation rerun of the original failed macOS job, with the exact head, baseline, workflow, and budgets unchanged. The failed sample above remains part of the evidence; no rerun-until-green loop or budget relaxation. Independent review accepted the control methodology and this single confirmation, but could not independently inspect the local raw-artifact directory because of an access boundary. The diagnostic artifact remains attached to the linked run for authorized review. This PR remains draft pending the confirmation and remaining gates.
Final quality gate at e79fcac: the single unchanged confirmation passed. macOS discovery P50/P95 are 144/180 ms, refresh RTT 145/181 ms, and cold RTT P50 296 ms; all 10 environment and 1 manager identities match. This is consistent with the same-host control's variability, though slower than the historical baseline. The original failing 205 ms sample and both attempts' artifacts remain preserved; no budgets, baseline, or runtime source were changed.
All 36 automated checks now pass. Current-head CodeQL analyses for Rust, Python, and Actions report 0 results and no analysis errors. Coverage is non-regressing on all three platforms and native subprocess attribution is proven. Current-head Copilot has no actionable findings; its pending-macOS-coverage note is superseded by the completed native run.
The accepted noisy metric is macOS host-to-host discovery/RTT variance, supported by the eight-pass exact-ref control rather than merely a green retry. I am marking this ready and enabling protected auto-merge. The remaining VS Code policy gate requires one collaborator approval (0/1); it will not be bypassed. The diagnostic control branch is not part of this PR and will not be merged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Make coverage evidence demonstrate real server execution and expose production gaps without changing existing whole-workspace regression gates.
Validation: 76 Python tests, 19 Windows native cases, workspace formatting/Clippy, and independent review pass. Actual Windows LLVM profiles prove all three execution witnesses increase from 0 to 1. Hosted macOS validation and quality inspection remain pending; no signing/release services run.
Fixes #534