Skip to content

workloads/eclair: instrument JARs at build time instead of with an agent - #276

Open
erickcestari wants to merge 1 commit into
lnfuzz:masterfrom
erickcestari:eclair-offline-instrumentation
Open

erickcestari wants to merge 1 commit into
lnfuzz:masterfrom
erickcestari:eclair-offline-instrumentation

Conversation

@erickcestari

Copy link
Copy Markdown
Member

The agent re-ran the ASM transform on every execution for classes first loaded after the snapshot. Rewriting the JARs at image build gives the same probes and IDs; bench-exec funding flow goes from 2.53 to 3.12 execs/s.

I'll also provide a coverage report for it.

The agent re-ran the ASM transform on every execution for classes first
loaded after the snapshot. Rewriting the JARs at image build gives the same
probes and IDs; bench-exec funding flow goes from 2.53 to 3.12 execs/s.
@erickcestari

erickcestari commented Oct 2, 2026 •

Copy link
Copy Markdown
Member Author

Fuzzing Evaluation Report

Configuration A (Baseline): master (Commit: 8d53f33 (2026-10-01))
Configuration B (Experimental): offline-inst (Commit: bac1c2a (2026-10-01))

1. Summary Statistics

Target Duration (h) n (Baseline) n (Exp.) Median Cov. (Baseline) Median Cov. (Exp.) Adj. p-value (Cov.) Â12 (Cov.) Median AUC (Baseline) Median AUC (Exp.) Adj. p-value (AUC) Â12 (AUC) Union Cov. (Baseline) Union Cov. (Exp.) Execs/s (Baseline) Execs/s (Exp.)
eclair 8 5 5 16238 16448 0.690476 0.6 127891 128818 0.420635 0.68 16914 16803 4.65 6.52

A comprehensive version of this table including raw P-values and Interquartile Ranges (IQRs) is available in evaluation_metrics.csv.

2. Interpretation Guide

Use the generated matrix above to objectively evaluate the experimental configuration. For full methodology, see the Smite Fuzzing Evaluation Framework.

Key Metrics

  • Adj. p-value: Mann-Whitney U test corrected for multiple targets via Holm-Bonferroni. Controls false-positive rate to ≤ 5% across all targets.
  • Â12: Probability that a random B trial outperforms a random A trial. 0.5 = no difference; 0.7 = B wins 70% of pairings. Always read alongside the p-value.
  • IQR: Spread of the middle 50% of trials. A much larger IQR in B suggests a few outlier runs may be inflating the median.
  • AUC: Coverage speed — how much was discovered and how early. Useful when final coverage is similar between configurations.
  • Union Coverage: Union of all trial bitmaps; the coverage ceiling for a multi-core deployment. Descriptive only, cannot be statistically tested.
  • Execs/s: A large drop in B without a coverage gain means the new feature is too expensive.

Reading the Results

Adj. p Â12 Conclusion
< 0.05 > 0.5 Meaningful improvement. Check IQRs are comparable, then merge.
< 0.05 ~0.5 Significant but negligible. Check if worth the added complexity.
> 0.05 > 0.6 Promising but underpowered. Re-run with more trials (e.g., 50).
> 0.05 ~0.5 No effect. Try an advanced snapshot or ground-truth evaluation.
any < 0.5 B underperforms A. If significant, reject or redesign the feature.

Time-series caveat: If the IQR bands overlap for most of the campaign and only diverge near the end, treat the final-coverage result cautiously — late divergence may reflect noise rather than a sustained advantage.

3. Visualizations

Note: In the box plots below, the central box represents the Interquartile Range (IQR, the middle 50% of trials), demonstrating the consistency of the fuzzer's performance. The internal line represents the median.

Target: eclair

Median Coverage Over Time

image

Distribution Comparisons

image

@erickcestari

Copy link
Copy Markdown
Member Author

The results are a bit noise. I will run another 6 trials of 24 hours each.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant