SD2 (System Deviation Diagnosis) is a robustness diagnosis framework for end-to-end (E2E) autonomous driving models. It decomposes a driving system into observable functional stages โ perception โ scene representation โ planning โ control โ outcome โ runs the same scenario under clean and stressed conditions, measures how much each stage deviates, and localizes the stage where robustness first collapses and how the error propagates downstream.
Instead of asking "how well does this model drive?", SD2 asks "where in the pipeline does robustness collapse, and how does the error propagate?"
Diagnosis outputs are temporal-correlational: SD2 identifies the earliest stage whose deviation crosses calibrated thresholds and is temporally followed by downstream deviation and/or driving-failure evidence. It does not claim a mechanistic root cause.
A modern E2E model is a black box from sensors to actuation. SD2 "opens" it into observable functional stages and reads the intermediate state at each one โ raw image, neural features, the model's internal scene representation (object density, BEV detections, or BEV occupancy), the predicted trajectory, and the vehicle control signals:
We deliberately say scene representation rather than scene understanding: InterFuser's object density and NEAT's BEV occupancy are semantic representations, not human-style understanding. A reasoning stage also exists in the schema, but it is optional and used only by language-based driving agents โ it plays no part in the E2E experiments reported here.
SD2 pairs a clean run with a stress run frame by frame, computes a normalized deviation per stage, analyzes how deviations propagate between adjacent stages, and diagnoses the primary failure stage โ then writes a Markdown report:
The figures and summary below come from recordings made before the 2026-07-10 pipeline audit and are not trustworthy. Any route-completion figure in them is an artifact of the progress-tracker defect. They are kept only to show the shape of SD2's output and will be regenerated once re-recording finishes.
A CARLA closed-loop run produces a per-stage robustness fingerprint โ higher is more robust:
The deviation timeline shows where robustness degrades first, and the diagnosis module turns it into a natural-language summary naming the primary failure stage, the downstream propagation order, and the driving outcome:
diagnosis.json includes "diagnosis_type": "temporal_correlational" to make this framing explicit โ SD2 localizes the earliest-collapsing stage by timing, not by mechanistic proof.
See the full generated report at docs/example/example_report.md.
Because SD2 is architecture-agnostic, it can diagnose different E2E models under
the same stress and reveal that they fail at different stages. Established on
re-recorded runs: InterFuser and NEAT fail at different stages under the same
corruption โ "divergent semantic collapse" โ including the safety-critical
RUSH probe red-light result. The same Gaussian noise freezes InterFuser
(phantom obstacles, semantic object-density stage) and blinds NEAT into running
red lights (semantic red_light_occ signal); restoring a single internal
semantic signal reverses each. See "Confirmed result: NEAT, divergent semantic
collapse, and the RUSH probe" below, with full numbers and scope in
docs/sd2_outcome_level_result.md.
Note: earlier cross-model claims in this section (from before the pipeline audit) were withdrawn; the results above are the re-recorded replacement.
SD2 diagnoses published E2E driving models; this study focuses on the two that
yield a localizable stressor-induced failure (InterFuser and NEAT). The weights and code are
consumed read-only through gitignored models/ junctions; nothing in this repo
redistributes them. Original sources:
| Model | Source repository | Paper (venue) | Sensors | Stages SD2 records | Semantic on the control path? |
|---|---|---|---|---|---|
| InterFuser | opendilab/InterFuser | Shao et al., CoRL 2022 | camera + LiDAR | vision, semantic (object density), planning, control | Yes โ object-density map feeds the controller (valid semantic intervention) |
| NEAT | autonomousvision/neat | Chitta et al., ICCV 2021 | multi-camera | vision, semantic (BEV occupancy + red_light_occ), planning, control |
Partly โ the recorded BEV-occupancy map is a side output off the path; only red_light_occ feeds the controller |
"Stages SD2 records" means SD2 logs a representation at that stage โ it does not imply that stage lies on the causal path to control. Whether the semantic representation actually feeds the controller (and is therefore a valid intervention target) is the last column, and it matches the intervention support matrix in the "Counterfactual stage intervention" section below. InterFuser exposes an on-path semantic signal (object density) everywhere; NEAT exposes one (red_light_occ) only where a red light is present, while its recorded BEV-occupancy map is off-path. The recorder is architecture-agnostic, but this study reports InterFuser and NEAT โ the two models that yield a localizable stressor-induced failure under 0.9.10-era checkpoints in CARLA 0.9.16.
Withdrawn pending re-recording (2026-07-10). An audit of the recording pipeline found nine defects, two of which corrupted the driving itself and the metric used to report it. Every live CARLA number previously published in this section โ robustness fingerprints, cross-stress and cross-town tables, and every route-completion figure โ was produced by that broken pipeline and has been removed rather than quietly restated.
The evidence, the isolation experiments, and the verification of each fix are in docs/recorder_bug_audit.md. The two defects that matter:
- The global plan was mirrored about y = 0.
_location_to_gpsused the CARLA 0.9.10 convention, but 0.9.16's GNSS reports latitude increasing with+y. Every model was steering toward a reflected goal. Measured live, the plan-vs-sensor delta was 48.99 m; after the fix, 0.16 m.- Route completion was fabricated by an index teleport. The progress tracker searched the whole remaining route for the nearest waypoint, so where a route passed near itself the monotonic index jumped (12 โ 254 in one frame) and completion leapt from 0.019 to 0.85. Replaying a real recording, a run reported as 80.8 % complete had covered 22 m of a 287 m route โ about 7.7 %.
Re-recording is under way. The methodology below is unchanged and still holds; only the measurements were invalid.
Different architectures expose different stages, so a single average is not comparable across models. SD2 therefore reports two summaries:
- Observed-stage mean โ averages whichever stages that model exposes. Use it for within-model diagnosis. It is not a cross-model ranking: a model is penalised simply for exposing a fragile stage that another model hides.
- Common-stage mean โ averages the stages every compared model exposes on
the causal path (e.g.
vision + planning + control), since models differ in which semantic signal reaches the controller. Use it for cross-model comparison.
sd2 fingerprint emits both columns, and fingerprint.json carries
common_stage_mean alongside mean_robustness.
SD2 includes a synthetic fault-injection benchmark for validating the diagnosis
framework itself. This is a framework sanity check, not a real-model experiment:
it creates clean/stress JSONL run pairs where the primary failure stage is known
by construction, runs the full run_analysis pipeline, and scores
diagnosis.json against the label.
The five labeled fault classes are:
vision: large visual embedding cosine deviationsemantic: object-set collapse plus critical-object and traffic-light flipsreasoning: intent/text/critical-object mention mismatchplanning: waypoint and target-speed divergencecontrol: steer/throttle/brake command spike
Run the benchmark with:
sd2 benchmark --config configs/mvp.yaml --output outputs/fault_benchmark --n-per-class 20 --seed 42Or run the demo wrapper:
python experiments/run_fault_benchmark.pyThe default demo writes benchmark_result.json, benchmark_report.md, and
confusion_matrix.png. The current example result is 100.0% overall accuracy
with 100.0% per-class accuracy on all five synthetic classes; this indicates
that the implemented diagnosis policy matches the controlled synthetic origins.
See docs/example/benchmark_report.md and
the embedded confusion heatmap:
For live CARLA recordings, SD2 also supports same-pose counterfactual stage
intervention. At each recorded tick the recorder reads the raw camera image as
I_clean, computes I_stress = stressor(I_clean), runs the model once on
I_stress and once on I_clean, then routes one selected stage from one
forward into the downstream controller. The two forwards occur at the same ego
pose, target point, and velocity, so recovery or failure is not explained by
clean and stress runs drifting into different closed-loop states before the
comparison.
Use:
python experiments/interfuser_record.py --stress gaussian_noise --stress-severity 3 --intervene-stage planning --intervene-direction restore --output data/carla/interfuser_restore_planning.jsonl
python experiments/interfuser_record.py --stress gaussian_noise --stress-severity 3 --intervene-stage semantic --intervene-direction restore --output data/carla/interfuser_restore_semantic.jsonl
sd2 intervention --baseline-clean data/carla/interfuser_clean.jsonl --stress data/carla/interfuser_stress.jsonl --intervened data/carla/interfuser_restore_planning.jsonl --config configs/mvp.yaml --output outputs/interfuser_restore_planning--intervene-direction restore means the run is a stress run, but the selected
stage is taken from the clean forward. It tests whether fixing that stage
recovers driving. --intervene-direction inject means the run is otherwise
clean, but the selected stage is taken from the stressed forward. It tests
whether breaking only that stage reproduces the failure. --intervene-stage none
still logs both forwards and both candidate controls while applying the
normal stressed control.
Support matrix:
| Model | planning |
semantic |
|---|---|---|
| InterFuser | supported | supported |
| NEAT | supported | supported (red_light_occ) |
Important caveat on control-level evidence: for a single-input controller that
consumes only planning waypoints and velocity, restoring planning at a fixed pose
restores the per-tick control by construction โ arithmetic, not empirical
mediation evidence โ so the informative evidence there is the closed-loop outcome
(route completion, collisions, lane invasions). The genuinely informative
control-level decomposition exists for multi-input controllers, i.e. both
models used here: InterFuser (object-density map + waypoints feed the controller)
and NEAT (red_light_occ + waypoints feed the controller).
On short, straight, spawn-aligned routes where InterFuser drives stably (clean route completion ~0.99, noise floor sd ~0.0005), the counterfactual chain closes at the closed-loop route-completion level, on two independent routes and across three analysis levels. Under gaussian_noise s5 InterFuser stalls completely; restoring the clean semantic stage recovers it, restoring planning does not:
| condition (route 31->36) | route completion |
|---|---|
| clean | 0.9990 ยฑ 0.0005 |
| gaussian_noise s5 | 0.0000 ยฑ 0.0000 |
| + semantic-restore | 0.9778 ยฑ 0.0001 (97.9% recovery) |
| + planning-restore | 0.0000 ยฑ 0.0000 (0%) |
Route 53->107 replicates (semantic recovery 99.1%, planning 1.0%). Different corruptions localize to different stages: gaussian_noise -> semantic (total stall), motion_blur -> planning (a 0.10 degradation with collisions, recovered only by planning-restore); brightness and fog do not meaningfully degrade this route. On route 31->36 the correlational diagnosis labels the stall "planning", but the counterfactual localizes it to "semantic" and the outcome recovery proves the counterfactual right โ the counterfactual is stable where the correlational label is not. Full protocol, numbers, and scope in docs/sd2_outcome_level_result.md. Other published E2E checkpoints we tried do not drive under 0.9.10-era weights in CARLA 0.9.16 (they full-brake or crawl), so they offer no localizable stressor-induced failure to analyze; NEAT does, and is covered next.
The same counterfactual, run on NEAT (a different architecture), localizes NEAT's failures to different stages than InterFuser โ and yields the sharpest, safety-critical result in the study.
- Contrast slowdown โ planning. A contrast_shift s5 corruption slows NEAT (~0.93 โ ~0.68-0.75 route completion on two routes); restoring the clean planning waypoints recovers it (necessity on both routes, bootstrap CIs exclude zero), while its semantic channel is scenario-vacuous on empty roads.
- Noise-induced red-light blindness โ semantic (
red_light_occ). At forced red lights, gaussian_noise s5 suppresses NEAT's internal red-light signalred_light_occfrom ~3.9 to 0; NEAT then drives through the red (4.7-33 m, deterministic across 6 seeds on 6 intersections). Restoring only that one clean signal โ every other perception channel still corrupted โ returns NEAT to a full stop at the line on 6/6 routes (mean 19.8 m of running eliminated, bootstrap 95% CIs far from zero); restoring the planning waypoints instead does not, on 4/6 routes.
| condition (forced-red route, disp in m; low = safe stop) | spawn 146 | spawn 21 |
|---|---|---|
| clean | 0.2 (stop) | 0.0 (stop) |
| gaussian_noise s5 (blind) | 33.5 (runs red) | 5.9 (runs red) |
+ red_light_occ-restore |
0.1 (stops) | 0.0 (stops) |
| + planning-restore | 33.3 (runs) | 5.7 (runs) |
This is divergent semantic collapse: the same corruption attacks the semantic-perception stage of both models but diverges into opposite catastrophic behaviours โ InterFuser freezes (phantom obstacles), NEAT runs reds (blindness) โ each reversible by restoring one internal semantic signal. We call that diagnostic the RUSH probe (Restore-a-Unit-Semantic-signal). Honest scope: the lights are forced red (controlled, not natural cycling); the red-light result holds on the 6 of 31 spawns where clean NEAT reliably stops at all (NEAT's OOD red-light stopping is otherwise unreliable); same 0.9.10-in-0.9.16 checkpoint caveat throughout. Details, CIs, and the naming note are in docs/sd2_outcome_level_result.md ยง10.
The benchmark also has a hard profile with competing and ambiguous faults:
competing collapse, strong propagation, near-simultaneous adjacent collapse,
and noisy upstream distractors. These samples are labeled by the intended
origin stage, and near-simultaneous cases carry ambiguous: true in
label.json. The hard tier is meant to be discriminative, so accuracy below
100% is expected and should be reported honestly.
sd2 benchmark --config configs/mvp.yaml --output outputs/fault_benchmark_hard --profile hard --n-per-class 20 --seed 42Hard reports keep the confusion matrix and add per-ambiguity-type accuracy plus an ambiguous-only accuracy slice.
The
reasoningstage is optional and applies only to language-based driving agents that emit text. None of the E2E models benchmarked above expose it, and it is not part of the coreperception โ scene representation โ planning โ control โ outcomepipeline. This section is retained for language-agent adapters.
The default reasoning metric remains text_embedding_and_intent with weights
text_embedding=0.5, intent_mismatch=0.3, and
critical_object_mismatch=0.2. Three ablation variants are registered for
review analysis: reasoning_intent_only, reasoning_text_only, and
reasoning_critical_object_only.
The current text_embedding component is a token-set Jaccard distance, not a
semantic embedding. It is intentionally documented as paraphrase-fragile:
same-meaning rewrites can score as large lexical deviations, while small word
edits can hide decision changes. Intent weighting mitigates this for the MVP,
and the ablation probe motivates a future embedding or judge-based upgrade.
See docs/example/reasoning_ablation.md.
Create and activate the conda environment:
conda create -n sd2 python=3.12 -y
conda activate sd2Install the package in editable mode:
pip install -e .Run the one-command MVP demo:
python experiments/run_mvp.pyOr run analysis and report generation directly:
python -m sd2.cli analyze --clean data/sample/clean_run.jsonl --stress data/sample/stress_run.jsonl --config configs/mvp.yaml --output outputs/sample_analysis --reportCalibrate warning/critical thresholds from repeated clean runs:
python -m sd2.cli calibrate --clean clean_a.jsonl --clean clean_b.jsonl --clean clean_c.jsonl --config configs/mvp.yaml --output outputs/calibrationConsume calibrated per-stage thresholds during analysis:
python -m sd2.cli analyze --clean data/sample/clean_run.jsonl --stress data/sample/stress_run.jsonl --config configs/mvp.yaml --thresholds outputs/calibration/calibrated_thresholds.json --output outputs/sample_analysis_calibrated --reportGenerate a report from an existing analysis directory:
python -m sd2.cli report --analysis-dir outputs/sample_analysisAggregate one or more fingerprint outputs:
python -m sd2.cli fingerprint --analysis-dir outputs --output outputs/fingerprint_summary.mdGenerate deterministic sample images and apply a stressor:
python experiments/generate_sample_images.py
python -m sd2.cli stress --input data/sample/images --config configs/stress/gaussian_noise.yaml --output outputs/stress_demo --seed 42Run tests:
conda run -n sd2 python -m pytest -qIf Windows temp permissions interfere with pytest, use:
conda run -n sd2 python -m pytest -q --basetemp .pytest_basetempClean and stress runs are paired by pairing.mode in the YAML config. The
default is frame_idx, which preserves the original MVP behavior: only equal
frame indices are paired, and the pair key remains
model_id:scenario_id:seed:<clean_frame_idx>.
Two alternate clean-centric anchors are available for closed-loop runs where
stress can change the ego trajectory. timestamp pairs each clean frame with
the nearest stress timestamp within pairing.timestamp_tolerance seconds.
route_progress pairs each clean frame with the nearest stress
outcome.route_progress within pairing.progress_tolerance; this is useful
when heavy stress makes the ego lag, drift, or reach intersections at different
frame numbers. Route-progress mode requires outcome.route_progress on both
runs; otherwise use frame_idx.
For every mode, emitted pairs keep the clean frame's frame_idx and
timestamp, so deviation timelines, propagation, and onset logic remain on the
clean-run timeline. pairing_summary.json reports mode,
mean_anchor_mismatch, and max_anchor_mismatch; units are frame-index delta
for frame_idx (always 0.0), seconds for timestamp, and route-progress
fraction for route_progress. This addresses the known caveat that pure
frame-index pairing can misalign comparable driving states under heavy stress.
CARLA is not a core package dependency because its wheel is local and
platform-specific. Install CARLA 0.9.16 manually inside the sd2 environment:
pip install external/Carla/CARLA_0.9.16/PythonAPI/carla/dist/carla-0.9.16-*.whlThe recording client drives with CARLA's BasicAgent, which requires two extra
packages:
pip install shapely networkxLaunch the CARLA server from the CARLA install directory:
CarlaUE4.exe -quality-level=Low -RenderOffScreen -carla-rpc-port=2000Record a clean run:
python experiments/carla_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 200 --warmup 20 --seed 42 --delta 0.05 --stress none --output data/carla/town10_clean_seed42.jsonl --spawn-index 0Record a matched control_noise stress run with the same seed, town, frame
count, and spawn index:
python experiments/carla_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 200 --warmup 20 --seed 42 --delta 0.05 --stress control_noise --stress-severity 3 --output data/carla/town10_control_noise_s3_seed42.jsonl --spawn-index 0Analyze the pair:
sd2 analyze --clean data/carla/town10_clean_seed42.jsonl --stress data/carla/town10_control_noise_s3_seed42.jsonl --config configs/mvp.yaml --output outputs/carla_control_noise_s3 --reportThe CARLA recorder currently populates only Planning, Control, and Outcome
states: waypoints, target speed, ego pose/speed, vehicle controls, collision,
lane invasion, route progress, and optional TTC. Vision, Semantic, and
Reasoning are intentionally absent until a model adapter is available, so these
logs are Observability Tier 0/1. Clean and stress runs are designed to pair by
frame_idx; severe weather or control noise can still make the closed-loop
trajectory diverge, so frame pairing is an alignment convention for analysis.
SD2 can record InterFuser, an E2E camera+lidar model, in CARLA and emit Tier 2/3 logs with Vision, Semantic, Planning, Control, and Outcome populated. The recorder is experiments/interfuser_record.py; the CARLA-free conversion module is src/sd2/adapters/interfuser_adapter.py.
models/InterFuser/ is expected to be a local junction to the InterFuser repo
and remains gitignored. The script applies the verified inference preamble
internally: it stubs imgaug, prepends models/InterFuser/interfuser so the
vendored timm 0.4.13 wins, and prepends the InterFuser leaderboard and
scenario_runner paths. Point --checkpoint at your own InterFuser weights,
e.g. via an environment variable:
export INTERFUSER_CKPT=/path/to/interfuser.pthThe recorder attaches the InterFuser sensor rig from the leaderboard agent:
front RGB 800x600 fov 100, left/right RGB 400x300 yaw -60/+60, lidar
ray_cast yaw -90, IMU, GNSS, and a speedometer measurement derived from the
ego velocity. Visual stressors are applied to RGB frames before InterFuser
preprocessing and inference, so the perturbation can propagate through
semantic prediction, planning, control, and outcome.
Record a clean InterFuser run:
python experiments/interfuser_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint "$INTERFUSER_CKPT" --stress none --output data/carla/interfuser_town10_clean_seed42.jsonl --spawn-index 0Record a matched Gaussian-noise stress run with the same seed, town, frame count, and spawn index:
python experiments/interfuser_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint "$INTERFUSER_CKPT" --stress gaussian_noise --stress-severity 3 --output data/carla/interfuser_town10_gaussian_noise_s3_seed42.jsonl --spawn-index 0Analyze the pair:
sd2 analyze --clean data/carla/interfuser_town10_clean_seed42.jsonl --stress data/carla/interfuser_town10_gaussian_noise_s3_seed42.jsonl --config configs/mvp.yaml --output outputs/interfuser_town10_gaussian_noise_s3 --reportStage mapping:
vision: mean-pooled InterFusertraffic_feature/BEV feature asfeatureforembedding_cosine, plus front-cameraimage_meanandimage_stdfallback.semantic: trackedtraffic_metaobject counts/classes, occupied-cell density, junction probability, traffic-light score, and stop-sign score.planning: predicted waypoints, controller target speed, route command, and local target point.control:InterfuserControllersteer, throttle, and brake.outcome: CARLA collision and lane-invasion events, route progress, and optional TTC placeholder.
The script logs first-tick sensor shapes, model input tensor shapes, and model output shapes before recording frames, which is the first place to look if a live CARLA run has an input-shape mismatch.
SD2 records NEAT as the second architecture, using the same clean/stress pairing and SD2 stage schema:
- NEAT: attention-field model; SD2 observes multi-camera encoder features,
decoded BEV occupancy semantics (
bev_seg_summary), predicted waypoints, PID control, and outcome. Thebev_seg_summarymap is adecode(...)side output off the control path; the only NEAT semantic signal that actually feedscontrol_pidisred_light_occ, so that โ not the BEV map โ is what a semantic intervention on NEAT swaps.
models/NEAT/ is expected to be a local gitignored junction. Default checkpoints:
models/NEAT/neat/best_encoder.pth
models/NEAT/neat/best_decoder.pth
models/NEAT/neat/args.txt
Over 300 frames with correctly configured sensors and no nudge, NEAT drives (it does not fall into the cold-start crawl that afflicts some other checkpoints), and its confirmed SD2 results are the contrastโplanning and red-lightโsemantic (RUSH) localizations summarized above and in docs/sd2_outcome_level_result.md ยง10.
Record and analyze NEAT:
python experiments/neat_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint models/NEAT/neat --stress none --output data/carla/neat_town10_clean_seed42.jsonl --spawn-index 0
python experiments/neat_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint models/NEAT/neat --stress gaussian_noise --stress-severity 3 --output data/carla/neat_town10_gaussian_noise_s3_seed42.jsonl --spawn-index 0
sd2 analyze --clean data/carla/neat_town10_clean_seed42.jsonl --stress data/carla/neat_town10_gaussian_noise_s3_seed42.jsonl --config configs/mvp.yaml --output outputs/neat_town10_gaussian_noise_s3 --reportAggregate fingerprints across recorded runs:
sd2 fingerprint --analysis-dir outputs --output outputs/e2e_fingerprint_summary.md
sd2 aggregate --analysis-dir outputs --output outputs/e2e_aggregate.mdReal CARLA drives have natural run-to-run variation (engine non-determinism), so repeated clean runs give a meaningful clean-clean baseline. Record several clean runs and calibrate per-stage thresholds, then analyze the stress pair with them:
sd2 calibrate --clean data/carla/calib/clean_rep1.jsonl --clean data/carla/calib/clean_rep2.jsonl --clean data/carla/calib/clean_rep3.jsonl --config configs/mvp.yaml --output data/carla/calib/calibrated_thresholds.json
sd2 analyze --clean data/carla/town10_clean_seed42.jsonl --stress data/carla/town10_control_noise_s3_seed42.jsonl --config configs/mvp.yaml --thresholds data/carla/calib/calibrated_thresholds.json --output outputs/carla_control_noise_s3_calibrated --reportOn the bundled control_noise example this matters: the static 0.4/0.7
thresholds are far above the real clean-clean deviation (Planning/Control
critical calibrate to about 0.06-0.07), and a single-frame Planning spike
would otherwise be mistaken for the primary failure. With calibrated thresholds
plus the onset_persistence_frames requirement (a collapse onset must persist
for several consecutive frames, filtering outlier spikes), the diagnosis
correctly identifies Control โ the stage that was actually perturbed โ as
the primary failure stage.
The demo writes:
outputs/sample_analysis/paired_frames.jsonoutputs/sample_analysis/pairing_summary.jsonoutputs/sample_analysis/deviation_table.jsonoutputs/sample_analysis/deviation_table.csvoutputs/sample_analysis/propagation.jsonoutputs/sample_analysis/diagnosis.jsonoutputs/sample_analysis/fingerprint.jsonoutputs/sample_analysis/report.mdoutputs/sample_analysis/plots/deviation_timeline.pngoutputs/sample_analysis/plots/robustness_fingerprint.pngoutputs/sample_analysis/plots/propagation_scores.png
Calibration writes calibrated_thresholds.json, containing per-stage clean-clean mean/std, warning/critical thresholds computed as mean + k * std, and fallback flags for stages whose clean-clean variance is near zero.
A copy of the demo output (report and plots) is kept under docs/example/ for reference.
Stressors perturb clean image inputs to produce offline stress-run inputs. Severity is an integer from 1 to 5; 0 and out-of-range values are rejected. Each stressor maps that severity to concrete parameters internally and records those parameters in stress_manifest.json.
Visual stressors operate on HxWx3 RGB uint8 images and write the same filenames to the output directory:
gaussian_noisemotion_blurfogbrightness_shiftcontrast_shiftjpeg_compressionlow_light
Temporal stressors operate on the sorted image list as a frame sequence:
frame_dropframe_delaycamera_blackoutlow_fps
For temporal materialization, dropped frames are omitted, camera blackouts are written as black images, and delayed frames hold earlier source images at the current output position. Run a stress pass with:
python -m sd2.cli stress --input <image-dir> --config <stress-yaml> --output <output-dir> --seed 42Existing stress configs live in configs/stress/, for example gaussian_noise.yaml, motion_blur.yaml, and frame_drop.yaml.
MVP Phase 1 through the offline stressor layer are complete:
- src-layout Python package scaffold
- Pydantic v2 run and frame schema
- deterministic JSONL sample data
- deterministic sample image generator for stressor demos
- JSONL run loader with line-numbered validation errors
- clean/stress frame pairing with skipped-frame summary and saved run metadata
- stage-wise metric registry and MVP metrics for vision, semantic, planning, and control stages (plus an optional reasoning stage for language-based agents)
- visual and temporal stressor registry with
sd2 stressCLI materialization - min-max clipping and threshold status classification (
healthy,warning,critical) - optional clean-clean threshold calibration with
sd2 calibrateandsd2 analyze --thresholds - propagation analysis with adjacent-stage robust evidence bundles: legacy ratio, clipped ratio, log-ratio, absolute increase, collapse order, and downstream persistence
- temporal-correlational failure-stage labeling using
first_critical_with_downstream_increase, with documented fallbacks - per-run robustness fingerprint where each observed stage score is
1 - mean(normalized deviation) - Markdown report generation with stage timeline, fingerprint, and propagation plots
sd2 analyze --report,sd2 report, andsd2 fingerprintCLI flowsexperiments/run_mvp.pyone-command demo- labeled synthetic fault-injection benchmark with
sd2 benchmark - hard/ambiguous synthetic benchmark profile with per-ambiguity reporting
- optional-stage reasoning metric ablations and paraphrase-robustness probe (language-based agents only; unused by the E2E experiments)
experiments/run_fault_benchmark.pyone-command validation demo- CARLA InterFuser and NEAT E2E recorders plus pure SD2 adapters for stage-wise diagnosis
- observed-stage and common-stage robustness means (
sd2 fingerprint) and multi-seed statistical robustness (sd2 aggregate)
The synthetic benchmark validates the SD2 diagnosis machinery on controlled offline logs; it does not replace real-model robustness experiments.
Live CARLA results are being re-recorded. The 2026-07-10 pipeline audit found nine recorder defects, two of which invalidated every previously published live measurement. The recorders and the diagnosis machinery are fixed and covered by tests; the experiments themselves have not yet been re-run.
Metrics are selected per stage in configs/mvp.yaml under metrics:
embedding_cosineforvision: cosine distance overembeddingorfeature.object_jaccardforsemantic: object-set Jaccard distance with missing/extra object, critical-object mismatch, and traffic-light mismatch details.text_embedding_and_intentforreasoning: weighted lexical token-set distance, intent mismatch, and critical-object mention mismatch.waypoint_adeforplanning: ADE over common waypoint prefix, with FDE and target-speed difference details.weighted_action_maeforcontrol: weighted absolute steer/throttle/brake error.





