4.6. Collect defensible samples
A p95 inside one report describes calls in that process. It is not a confidence interval over independent runs. For a change that matters:
-
pin the Lean toolchain and Lake manifests;
-
use the same build settings and machine;
-
fix workload, seed, backend, dtype, device, and synchronization policy;
-
isolate warmup from active samples;
-
collect several independent baseline and candidate runs;
-
inspect the distribution before setting a tolerance;
-
retain traces and summaries for failed comparisons.
Never share a baseline between an unsynchronized GPU launch span and a synchronized completion span. They measure different intervals.