Skip to content
justinbrianhwangPublic

About

๐’ฎ๐’Ÿ-2 ยท System Deviation Diagnosis โ€” a robustness diagnosis framework for end-to-end (E2E) autonomous driving. Decomposes the driving pipeline (vision โ†’ semantic โ†’ planning โ†’ control โ†’ outcome), measures stage-wise deviation between clean and stress CARLA runs, and localizes where robustness first collapses (InterFuser, TransFuser).

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

ย 

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

SD2

SD2 (System Deviation Diagnosis) is a robustness diagnosis framework for end-to-end (E2E) autonomous driving models. It decomposes a driving system into observable functional stages โ€” perception โ†’ scene representation โ†’ planning โ†’ control โ†’ outcome โ€” runs the same scenario under clean and stressed conditions, measures how much each stage deviates, and localizes the stage where robustness first collapses and how the error propagates downstream.

Instead of asking "how well does this model drive?", SD2 asks "where in the pipeline does robustness collapse, and how does the error propagate?"

SD2 concept: where does robustness collapse in the driving pipeline

Diagnosis outputs are temporal-correlational: SD2 identifies the earliest stage whose deviation crosses calibrated thresholds and is temporally followed by downstream deviation and/or driving-failure evidence. It does not claim a mechanistic root cause.

What SD2 observes

A modern E2E model is a black box from sensors to actuation. SD2 "opens" it into observable functional stages and reads the intermediate state at each one โ€” raw image, neural features, the model's internal scene representation (object density, BEV detections, or BEV occupancy), the predicted trajectory, and the vehicle control signals:

Opening the E2E black box into functional stages

We deliberately say scene representation rather than scene understanding: InterFuser's object density and NEAT's BEV occupancy are semantic representations, not human-style understanding. A reasoning stage also exists in the schema, but it is optional and used only by language-based driving agents โ€” it plays no part in the E2E experiments reported here.

How it works

SD2 pairs a clean run with a stress run frame by frame, computes a normalized deviation per stage, analyzes how deviations propagate between adjacent stages, and diagnoses the primary failure stage โ€” then writes a Markdown report:

SD2 method flow: clean/stress runs to diagnosis report

Example Output

The figures and summary below come from recordings made before the 2026-07-10 pipeline audit and are not trustworthy. Any route-completion figure in them is an artifact of the progress-tracker defect. They are kept only to show the shape of SD2's output and will be regenerated once re-recording finishes.

A CARLA closed-loop run produces a per-stage robustness fingerprint โ€” higher is more robust:

Robustness fingerprint

The deviation timeline shows where robustness degrades first, and the diagnosis module turns it into a natural-language summary naming the primary failure stage, the downstream propagation order, and the driving outcome:

Stage-wise deviation timeline

diagnosis.json includes "diagnosis_type": "temporal_correlational" to make this framing explicit โ€” SD2 localizes the earliest-collapsing stage by timing, not by mechanistic proof.

See the full generated report at docs/example/example_report.md.

Cross-architecture comparison

Because SD2 is architecture-agnostic, it can diagnose different E2E models under the same stress and reveal that they fail at different stages. Established on re-recorded runs: InterFuser and NEAT fail at different stages under the same corruption โ€” "divergent semantic collapse" โ€” including the safety-critical RUSH probe red-light result. The same Gaussian noise freezes InterFuser (phantom obstacles, semantic object-density stage) and blinds NEAT into running red lights (semantic red_light_occ signal); restoring a single internal semantic signal reverses each. See "Confirmed result: NEAT, divergent semantic collapse, and the RUSH probe" below, with full numbers and scope in docs/sd2_outcome_level_result.md.

Note: earlier cross-model claims in this section (from before the pipeline audit) were withdrawn; the results above are the re-recorded replacement.

E2E models & source repositories

SD2 diagnoses published E2E driving models; this study focuses on the two that yield a localizable stressor-induced failure (InterFuser and NEAT). The weights and code are consumed read-only through gitignored models/ junctions; nothing in this repo redistributes them. Original sources:

Model Source repository Paper (venue) Sensors Stages SD2 records Semantic on the control path?
InterFuser opendilab/InterFuser Shao et al., CoRL 2022 camera + LiDAR vision, semantic (object density), planning, control Yes โ€” object-density map feeds the controller (valid semantic intervention)
NEAT autonomousvision/neat Chitta et al., ICCV 2021 multi-camera vision, semantic (BEV occupancy + red_light_occ), planning, control Partly โ€” the recorded BEV-occupancy map is a side output off the path; only red_light_occ feeds the controller

"Stages SD2 records" means SD2 logs a representation at that stage โ€” it does not imply that stage lies on the causal path to control. Whether the semantic representation actually feeds the controller (and is therefore a valid intervention target) is the last column, and it matches the intervention support matrix in the "Counterfactual stage intervention" section below. InterFuser exposes an on-path semantic signal (object density) everywhere; NEAT exposes one (red_light_occ) only where a red light is present, while its recorded BEV-occupancy map is off-path. The recorder is architecture-agnostic, but this study reports InterFuser and NEAT โ€” the two models that yield a localizable stressor-induced failure under 0.9.10-era checkpoints in CARLA 0.9.16.

Results: live CARLA robustness diagnosis

Withdrawn pending re-recording (2026-07-10). An audit of the recording pipeline found nine defects, two of which corrupted the driving itself and the metric used to report it. Every live CARLA number previously published in this section โ€” robustness fingerprints, cross-stress and cross-town tables, and every route-completion figure โ€” was produced by that broken pipeline and has been removed rather than quietly restated.

The evidence, the isolation experiments, and the verification of each fix are in docs/recorder_bug_audit.md. The two defects that matter:

  1. The global plan was mirrored about y = 0. _location_to_gps used the CARLA 0.9.10 convention, but 0.9.16's GNSS reports latitude increasing with +y. Every model was steering toward a reflected goal. Measured live, the plan-vs-sensor delta was 48.99 m; after the fix, 0.16 m.
  2. Route completion was fabricated by an index teleport. The progress tracker searched the whole remaining route for the nearest waypoint, so where a route passed near itself the monotonic index jumped (12 โ†’ 254 in one frame) and completion leapt from 0.019 to 0.85. Replaying a real recording, a run reported as 80.8 % complete had covered 22 m of a 287 m route โ€” about 7.7 %.

Re-recording is under way. The methodology below is unchanged and still holds; only the measurements were invalid.

Methodology that survives the audit

Two means, and when to use which

Different architectures expose different stages, so a single average is not comparable across models. SD2 therefore reports two summaries:

  • Observed-stage mean โ€” averages whichever stages that model exposes. Use it for within-model diagnosis. It is not a cross-model ranking: a model is penalised simply for exposing a fragile stage that another model hides.
  • Common-stage mean โ€” averages the stages every compared model exposes on the causal path (e.g. vision + planning + control), since models differ in which semantic signal reaches the controller. Use it for cross-model comparison.

sd2 fingerprint emits both columns, and fingerprint.json carries common_stage_mean alongside mean_robustness.

Validation: Synthetic Fault Injection Benchmark

SD2 includes a synthetic fault-injection benchmark for validating the diagnosis framework itself. This is a framework sanity check, not a real-model experiment: it creates clean/stress JSONL run pairs where the primary failure stage is known by construction, runs the full run_analysis pipeline, and scores diagnosis.json against the label.

The five labeled fault classes are:

  • vision: large visual embedding cosine deviation
  • semantic: object-set collapse plus critical-object and traffic-light flips
  • reasoning: intent/text/critical-object mention mismatch
  • planning: waypoint and target-speed divergence
  • control: steer/throttle/brake command spike

Run the benchmark with:

sd2 benchmark --config configs/mvp.yaml --output outputs/fault_benchmark --n-per-class 20 --seed 42

Or run the demo wrapper:

python experiments/run_fault_benchmark.py

The default demo writes benchmark_result.json, benchmark_report.md, and confusion_matrix.png. The current example result is 100.0% overall accuracy with 100.0% per-class accuracy on all five synthetic classes; this indicates that the implemented diagnosis policy matches the controlled synthetic origins. See docs/example/benchmark_report.md and the embedded confusion heatmap:

Synthetic benchmark confusion matrix

Validating the diagnosis: counterfactual intervention

For live CARLA recordings, SD2 also supports same-pose counterfactual stage intervention. At each recorded tick the recorder reads the raw camera image as I_clean, computes I_stress = stressor(I_clean), runs the model once on I_stress and once on I_clean, then routes one selected stage from one forward into the downstream controller. The two forwards occur at the same ego pose, target point, and velocity, so recovery or failure is not explained by clean and stress runs drifting into different closed-loop states before the comparison.

Use:

python experiments/interfuser_record.py --stress gaussian_noise --stress-severity 3 --intervene-stage planning --intervene-direction restore --output data/carla/interfuser_restore_planning.jsonl
python experiments/interfuser_record.py --stress gaussian_noise --stress-severity 3 --intervene-stage semantic --intervene-direction restore --output data/carla/interfuser_restore_semantic.jsonl
sd2 intervention --baseline-clean data/carla/interfuser_clean.jsonl --stress data/carla/interfuser_stress.jsonl --intervened data/carla/interfuser_restore_planning.jsonl --config configs/mvp.yaml --output outputs/interfuser_restore_planning

--intervene-direction restore means the run is a stress run, but the selected stage is taken from the clean forward. It tests whether fixing that stage recovers driving. --intervene-direction inject means the run is otherwise clean, but the selected stage is taken from the stressed forward. It tests whether breaking only that stage reproduces the failure. --intervene-stage none still logs both forwards and both candidate controls while applying the normal stressed control.

Support matrix:

Model planning semantic
InterFuser supported supported
NEAT supported supported (red_light_occ)

Important caveat on control-level evidence: for a single-input controller that consumes only planning waypoints and velocity, restoring planning at a fixed pose restores the per-tick control by construction โ€” arithmetic, not empirical mediation evidence โ€” so the informative evidence there is the closed-loop outcome (route completion, collisions, lane invasions). The genuinely informative control-level decomposition exists for multi-input controllers, i.e. both models used here: InterFuser (object-density map + waypoints feed the controller) and NEAT (red_light_occ + waypoints feed the controller).

Confirmed result: the causal chain closes at the outcome level

On short, straight, spawn-aligned routes where InterFuser drives stably (clean route completion ~0.99, noise floor sd ~0.0005), the counterfactual chain closes at the closed-loop route-completion level, on two independent routes and across three analysis levels. Under gaussian_noise s5 InterFuser stalls completely; restoring the clean semantic stage recovers it, restoring planning does not:

condition (route 31->36) route completion
clean 0.9990 ยฑ 0.0005
gaussian_noise s5 0.0000 ยฑ 0.0000
+ semantic-restore 0.9778 ยฑ 0.0001 (97.9% recovery)
+ planning-restore 0.0000 ยฑ 0.0000 (0%)

Route 53->107 replicates (semantic recovery 99.1%, planning 1.0%). Different corruptions localize to different stages: gaussian_noise -> semantic (total stall), motion_blur -> planning (a 0.10 degradation with collisions, recovered only by planning-restore); brightness and fog do not meaningfully degrade this route. On route 31->36 the correlational diagnosis labels the stall "planning", but the counterfactual localizes it to "semantic" and the outcome recovery proves the counterfactual right โ€” the counterfactual is stable where the correlational label is not. Full protocol, numbers, and scope in docs/sd2_outcome_level_result.md. Other published E2E checkpoints we tried do not drive under 0.9.10-era weights in CARLA 0.9.16 (they full-brake or crawl), so they offer no localizable stressor-induced failure to analyze; NEAT does, and is covered next.

Confirmed result: NEAT, divergent semantic collapse, and the RUSH probe

The same counterfactual, run on NEAT (a different architecture), localizes NEAT's failures to different stages than InterFuser โ€” and yields the sharpest, safety-critical result in the study.

  • Contrast slowdown โ†’ planning. A contrast_shift s5 corruption slows NEAT (~0.93 โ†’ ~0.68-0.75 route completion on two routes); restoring the clean planning waypoints recovers it (necessity on both routes, bootstrap CIs exclude zero), while its semantic channel is scenario-vacuous on empty roads.
  • Noise-induced red-light blindness โ†’ semantic (red_light_occ). At forced red lights, gaussian_noise s5 suppresses NEAT's internal red-light signal red_light_occ from ~3.9 to 0; NEAT then drives through the red (4.7-33 m, deterministic across 6 seeds on 6 intersections). Restoring only that one clean signal โ€” every other perception channel still corrupted โ€” returns NEAT to a full stop at the line on 6/6 routes (mean 19.8 m of running eliminated, bootstrap 95% CIs far from zero); restoring the planning waypoints instead does not, on 4/6 routes.
condition (forced-red route, disp in m; low = safe stop) spawn 146 spawn 21
clean 0.2 (stop) 0.0 (stop)
gaussian_noise s5 (blind) 33.5 (runs red) 5.9 (runs red)
+ red_light_occ-restore 0.1 (stops) 0.0 (stops)
+ planning-restore 33.3 (runs) 5.7 (runs)

This is divergent semantic collapse: the same corruption attacks the semantic-perception stage of both models but diverges into opposite catastrophic behaviours โ€” InterFuser freezes (phantom obstacles), NEAT runs reds (blindness) โ€” each reversible by restoring one internal semantic signal. We call that diagnostic the RUSH probe (Restore-a-Unit-Semantic-signal). Honest scope: the lights are forced red (controlled, not natural cycling); the red-light result holds on the 6 of 31 spawns where clean NEAT reliably stops at all (NEAT's OOD red-light stopping is otherwise unreliable); same 0.9.10-in-0.9.16 checkpoint caveat throughout. Details, CIs, and the naming note are in docs/sd2_outcome_level_result.md ยง10.

Hard Benchmark Tier

The benchmark also has a hard profile with competing and ambiguous faults: competing collapse, strong propagation, near-simultaneous adjacent collapse, and noisy upstream distractors. These samples are labeled by the intended origin stage, and near-simultaneous cases carry ambiguous: true in label.json. The hard tier is meant to be discriminative, so accuracy below 100% is expected and should be reported honestly.

sd2 benchmark --config configs/mvp.yaml --output outputs/fault_benchmark_hard --profile hard --n-per-class 20 --seed 42

Hard reports keep the confusion matrix and add per-ambiguity-type accuracy plus an ambiguous-only accuracy slice.

Reasoning Metric (optional stage): Ablations and Known Limitations

The reasoning stage is optional and applies only to language-based driving agents that emit text. None of the E2E models benchmarked above expose it, and it is not part of the core perception โ†’ scene representation โ†’ planning โ†’ control โ†’ outcome pipeline. This section is retained for language-agent adapters.

The default reasoning metric remains text_embedding_and_intent with weights text_embedding=0.5, intent_mismatch=0.3, and critical_object_mismatch=0.2. Three ablation variants are registered for review analysis: reasoning_intent_only, reasoning_text_only, and reasoning_critical_object_only.

The current text_embedding component is a token-set Jaccard distance, not a semantic embedding. It is intentionally documented as paraphrase-fragile: same-meaning rewrites can score as large lexical deviations, while small word edits can hide decision changes. Intent weighting mitigates this for the MVP, and the ablation probe motivates a future embedding or judge-based upgrade. See docs/example/reasoning_ablation.md.

Quickstart

Create and activate the conda environment:

conda create -n sd2 python=3.12 -y
conda activate sd2

Install the package in editable mode:

pip install -e .

Run the one-command MVP demo:

python experiments/run_mvp.py

Or run analysis and report generation directly:

python -m sd2.cli analyze --clean data/sample/clean_run.jsonl --stress data/sample/stress_run.jsonl --config configs/mvp.yaml --output outputs/sample_analysis --report

Calibrate warning/critical thresholds from repeated clean runs:

python -m sd2.cli calibrate --clean clean_a.jsonl --clean clean_b.jsonl --clean clean_c.jsonl --config configs/mvp.yaml --output outputs/calibration

Consume calibrated per-stage thresholds during analysis:

python -m sd2.cli analyze --clean data/sample/clean_run.jsonl --stress data/sample/stress_run.jsonl --config configs/mvp.yaml --thresholds outputs/calibration/calibrated_thresholds.json --output outputs/sample_analysis_calibrated --report

Generate a report from an existing analysis directory:

python -m sd2.cli report --analysis-dir outputs/sample_analysis

Aggregate one or more fingerprint outputs:

python -m sd2.cli fingerprint --analysis-dir outputs --output outputs/fingerprint_summary.md

Generate deterministic sample images and apply a stressor:

python experiments/generate_sample_images.py
python -m sd2.cli stress --input data/sample/images --config configs/stress/gaussian_noise.yaml --output outputs/stress_demo --seed 42

Run tests:

conda run -n sd2 python -m pytest -q

If Windows temp permissions interfere with pytest, use:

conda run -n sd2 python -m pytest -q --basetemp .pytest_basetemp

Pairing Anchors

Clean and stress runs are paired by pairing.mode in the YAML config. The default is frame_idx, which preserves the original MVP behavior: only equal frame indices are paired, and the pair key remains model_id:scenario_id:seed:<clean_frame_idx>.

Two alternate clean-centric anchors are available for closed-loop runs where stress can change the ego trajectory. timestamp pairs each clean frame with the nearest stress timestamp within pairing.timestamp_tolerance seconds. route_progress pairs each clean frame with the nearest stress outcome.route_progress within pairing.progress_tolerance; this is useful when heavy stress makes the ego lag, drift, or reach intersections at different frame numbers. Route-progress mode requires outcome.route_progress on both runs; otherwise use frame_idx.

For every mode, emitted pairs keep the clean frame's frame_idx and timestamp, so deviation timelines, propagation, and onset logic remain on the clean-run timeline. pairing_summary.json reports mode, mean_anchor_mismatch, and max_anchor_mismatch; units are frame-index delta for frame_idx (always 0.0), seconds for timestamp, and route-progress fraction for route_progress. This addresses the known caveat that pure frame-index pairing can misalign comparable driving states under heavy stress.

CARLA Logging (Real Closed-Loop)

CARLA is not a core package dependency because its wheel is local and platform-specific. Install CARLA 0.9.16 manually inside the sd2 environment:

pip install external/Carla/CARLA_0.9.16/PythonAPI/carla/dist/carla-0.9.16-*.whl

The recording client drives with CARLA's BasicAgent, which requires two extra packages:

pip install shapely networkx

Launch the CARLA server from the CARLA install directory:

CarlaUE4.exe -quality-level=Low -RenderOffScreen -carla-rpc-port=2000

Record a clean run:

python experiments/carla_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 200 --warmup 20 --seed 42 --delta 0.05 --stress none --output data/carla/town10_clean_seed42.jsonl --spawn-index 0

Record a matched control_noise stress run with the same seed, town, frame count, and spawn index:

python experiments/carla_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 200 --warmup 20 --seed 42 --delta 0.05 --stress control_noise --stress-severity 3 --output data/carla/town10_control_noise_s3_seed42.jsonl --spawn-index 0

Analyze the pair:

sd2 analyze --clean data/carla/town10_clean_seed42.jsonl --stress data/carla/town10_control_noise_s3_seed42.jsonl --config configs/mvp.yaml --output outputs/carla_control_noise_s3 --report

The CARLA recorder currently populates only Planning, Control, and Outcome states: waypoints, target speed, ego pose/speed, vehicle controls, collision, lane invasion, route progress, and optional TTC. Vision, Semantic, and Reasoning are intentionally absent until a model adapter is available, so these logs are Observability Tier 0/1. Clean and stress runs are designed to pair by frame_idx; severe weather or control noise can still make the closed-loop trajectory diverge, so frame pairing is an alignment convention for analysis.

E2E Model Diagnosis (InterFuser and NEAT)

SD2 can record InterFuser, an E2E camera+lidar model, in CARLA and emit Tier 2/3 logs with Vision, Semantic, Planning, Control, and Outcome populated. The recorder is experiments/interfuser_record.py; the CARLA-free conversion module is src/sd2/adapters/interfuser_adapter.py.

models/InterFuser/ is expected to be a local junction to the InterFuser repo and remains gitignored. The script applies the verified inference preamble internally: it stubs imgaug, prepends models/InterFuser/interfuser so the vendored timm 0.4.13 wins, and prepends the InterFuser leaderboard and scenario_runner paths. Point --checkpoint at your own InterFuser weights, e.g. via an environment variable:

export INTERFUSER_CKPT=/path/to/interfuser.pth

The recorder attaches the InterFuser sensor rig from the leaderboard agent: front RGB 800x600 fov 100, left/right RGB 400x300 yaw -60/+60, lidar ray_cast yaw -90, IMU, GNSS, and a speedometer measurement derived from the ego velocity. Visual stressors are applied to RGB frames before InterFuser preprocessing and inference, so the perturbation can propagate through semantic prediction, planning, control, and outcome.

Record a clean InterFuser run:

python experiments/interfuser_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint "$INTERFUSER_CKPT" --stress none --output data/carla/interfuser_town10_clean_seed42.jsonl --spawn-index 0

Record a matched Gaussian-noise stress run with the same seed, town, frame count, and spawn index:

python experiments/interfuser_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint "$INTERFUSER_CKPT" --stress gaussian_noise --stress-severity 3 --output data/carla/interfuser_town10_gaussian_noise_s3_seed42.jsonl --spawn-index 0

Analyze the pair:

sd2 analyze --clean data/carla/interfuser_town10_clean_seed42.jsonl --stress data/carla/interfuser_town10_gaussian_noise_s3_seed42.jsonl --config configs/mvp.yaml --output outputs/interfuser_town10_gaussian_noise_s3 --report

Stage mapping:

  • vision: mean-pooled InterFuser traffic_feature/BEV feature as feature for embedding_cosine, plus front-camera image_mean and image_std fallback.
  • semantic: tracked traffic_meta object counts/classes, occupied-cell density, junction probability, traffic-light score, and stop-sign score.
  • planning: predicted waypoints, controller target speed, route command, and local target point.
  • control: InterfuserController steer, throttle, and brake.
  • outcome: CARLA collision and lane-invasion events, route progress, and optional TTC placeholder.

The script logs first-tick sensor shapes, model input tensor shapes, and model output shapes before recording frames, which is the first place to look if a live CARLA run has an input-shape mismatch.

NEAT (second architecture)

SD2 records NEAT as the second architecture, using the same clean/stress pairing and SD2 stage schema:

  • NEAT: attention-field model; SD2 observes multi-camera encoder features, decoded BEV occupancy semantics (bev_seg_summary), predicted waypoints, PID control, and outcome. The bev_seg_summary map is a decode(...) side output off the control path; the only NEAT semantic signal that actually feeds control_pid is red_light_occ, so that โ€” not the BEV map โ€” is what a semantic intervention on NEAT swaps.

models/NEAT/ is expected to be a local gitignored junction. Default checkpoints:

models/NEAT/neat/best_encoder.pth
models/NEAT/neat/best_decoder.pth
models/NEAT/neat/args.txt

Over 300 frames with correctly configured sensors and no nudge, NEAT drives (it does not fall into the cold-start crawl that afflicts some other checkpoints), and its confirmed SD2 results are the contrastโ†’planning and red-lightโ†’semantic (RUSH) localizations summarized above and in docs/sd2_outcome_level_result.md ยง10.

Record and analyze NEAT:

python experiments/neat_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint models/NEAT/neat --stress none --output data/carla/neat_town10_clean_seed42.jsonl --spawn-index 0
python experiments/neat_record.py --host localhost --port 2000 --town Town10HD_Opt --frames 300 --warmup 20 --seed 42 --delta 0.05 --checkpoint models/NEAT/neat --stress gaussian_noise --stress-severity 3 --output data/carla/neat_town10_gaussian_noise_s3_seed42.jsonl --spawn-index 0
sd2 analyze --clean data/carla/neat_town10_clean_seed42.jsonl --stress data/carla/neat_town10_gaussian_noise_s3_seed42.jsonl --config configs/mvp.yaml --output outputs/neat_town10_gaussian_noise_s3 --report

Aggregate fingerprints across recorded runs:

sd2 fingerprint --analysis-dir outputs --output outputs/e2e_fingerprint_summary.md
sd2 aggregate --analysis-dir outputs --output outputs/e2e_aggregate.md

Calibrated thresholds on real data

Real CARLA drives have natural run-to-run variation (engine non-determinism), so repeated clean runs give a meaningful clean-clean baseline. Record several clean runs and calibrate per-stage thresholds, then analyze the stress pair with them:

sd2 calibrate --clean data/carla/calib/clean_rep1.jsonl --clean data/carla/calib/clean_rep2.jsonl --clean data/carla/calib/clean_rep3.jsonl --config configs/mvp.yaml --output data/carla/calib/calibrated_thresholds.json
sd2 analyze --clean data/carla/town10_clean_seed42.jsonl --stress data/carla/town10_control_noise_s3_seed42.jsonl --config configs/mvp.yaml --thresholds data/carla/calib/calibrated_thresholds.json --output outputs/carla_control_noise_s3_calibrated --report

On the bundled control_noise example this matters: the static 0.4/0.7 thresholds are far above the real clean-clean deviation (Planning/Control critical calibrate to about 0.06-0.07), and a single-frame Planning spike would otherwise be mistaken for the primary failure. With calibrated thresholds plus the onset_persistence_frames requirement (a collapse onset must persist for several consecutive frames, filtering outlier spikes), the diagnosis correctly identifies Control โ€” the stage that was actually perturbed โ€” as the primary failure stage.

Expected Outputs

The demo writes:

  • outputs/sample_analysis/paired_frames.json
  • outputs/sample_analysis/pairing_summary.json
  • outputs/sample_analysis/deviation_table.json
  • outputs/sample_analysis/deviation_table.csv
  • outputs/sample_analysis/propagation.json
  • outputs/sample_analysis/diagnosis.json
  • outputs/sample_analysis/fingerprint.json
  • outputs/sample_analysis/report.md
  • outputs/sample_analysis/plots/deviation_timeline.png
  • outputs/sample_analysis/plots/robustness_fingerprint.png
  • outputs/sample_analysis/plots/propagation_scores.png

Calibration writes calibrated_thresholds.json, containing per-stage clean-clean mean/std, warning/critical thresholds computed as mean + k * std, and fallback flags for stages whose clean-clean variance is near zero.

A copy of the demo output (report and plots) is kept under docs/example/ for reference.

Stressors

Stressors perturb clean image inputs to produce offline stress-run inputs. Severity is an integer from 1 to 5; 0 and out-of-range values are rejected. Each stressor maps that severity to concrete parameters internally and records those parameters in stress_manifest.json.

Visual stressors operate on HxWx3 RGB uint8 images and write the same filenames to the output directory:

  • gaussian_noise
  • motion_blur
  • fog
  • brightness_shift
  • contrast_shift
  • jpeg_compression
  • low_light

Temporal stressors operate on the sorted image list as a frame sequence:

  • frame_drop
  • frame_delay
  • camera_blackout
  • low_fps

For temporal materialization, dropped frames are omitted, camera blackouts are written as black images, and delayed frames hold earlier source images at the current output position. Run a stress pass with:

python -m sd2.cli stress --input <image-dir> --config <stress-yaml> --output <output-dir> --seed 42

Existing stress configs live in configs/stress/, for example gaussian_noise.yaml, motion_blur.yaml, and frame_drop.yaml.

Current Status

MVP Phase 1 through the offline stressor layer are complete:

  • src-layout Python package scaffold
  • Pydantic v2 run and frame schema
  • deterministic JSONL sample data
  • deterministic sample image generator for stressor demos
  • JSONL run loader with line-numbered validation errors
  • clean/stress frame pairing with skipped-frame summary and saved run metadata
  • stage-wise metric registry and MVP metrics for vision, semantic, planning, and control stages (plus an optional reasoning stage for language-based agents)
  • visual and temporal stressor registry with sd2 stress CLI materialization
  • min-max clipping and threshold status classification (healthy, warning, critical)
  • optional clean-clean threshold calibration with sd2 calibrate and sd2 analyze --thresholds
  • propagation analysis with adjacent-stage robust evidence bundles: legacy ratio, clipped ratio, log-ratio, absolute increase, collapse order, and downstream persistence
  • temporal-correlational failure-stage labeling using first_critical_with_downstream_increase, with documented fallbacks
  • per-run robustness fingerprint where each observed stage score is 1 - mean(normalized deviation)
  • Markdown report generation with stage timeline, fingerprint, and propagation plots
  • sd2 analyze --report, sd2 report, and sd2 fingerprint CLI flows
  • experiments/run_mvp.py one-command demo
  • labeled synthetic fault-injection benchmark with sd2 benchmark
  • hard/ambiguous synthetic benchmark profile with per-ambiguity reporting
  • optional-stage reasoning metric ablations and paraphrase-robustness probe (language-based agents only; unused by the E2E experiments)
  • experiments/run_fault_benchmark.py one-command validation demo
  • CARLA InterFuser and NEAT E2E recorders plus pure SD2 adapters for stage-wise diagnosis
  • observed-stage and common-stage robustness means (sd2 fingerprint) and multi-seed statistical robustness (sd2 aggregate)

The synthetic benchmark validates the SD2 diagnosis machinery on controlled offline logs; it does not replace real-model robustness experiments.

Live CARLA results are being re-recorded. The 2026-07-10 pipeline audit found nine recorder defects, two of which invalidated every previously published live measurement. The recorders and the diagnosis machinery are fixed and covered by tests; the experiments themselves have not yet been re-run.

Metric Config

Metrics are selected per stage in configs/mvp.yaml under metrics:

  • embedding_cosine for vision: cosine distance over embedding or feature.
  • object_jaccard for semantic: object-set Jaccard distance with missing/extra object, critical-object mismatch, and traffic-light mismatch details.
  • text_embedding_and_intent for reasoning: weighted lexical token-set distance, intent mismatch, and critical-object mention mismatch.
  • waypoint_ade for planning: ADE over common waypoint prefix, with FDE and target-speed difference details.
  • weighted_action_mae for control: weighted absolute steer/throttle/brake error.

About

๐’ฎ๐’Ÿ-2 ยท System Deviation Diagnosis โ€” a robustness diagnosis framework for end-to-end (E2E) autonomous driving. Decomposes the driving pipeline (vision โ†’ semantic โ†’ planning โ†’ control โ†’ outcome), measures stage-wise deviation between clean and stress CARLA runs, and localizes where robustness first collapses (InterFuser, TransFuser).

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages