Every failed model is trying
to tell you something.
Upload your evaluation data. WhyLab investigates why your model is failing, proves the root cause with counterfactual experiments, and repairs your operating policy.
- 01Observe
the failure signal - 02Investigate
with measured evidence - 03Test
the strongest hypothesis - 04Repair
and re-evaluate
High accuracy. Missed malignant cases.
A synthetic melanoma classifier exposes the accuracy paradox. Follow the evidence, test competing explanations, then inspect a measured policy repair.
Synthetic teaching data: 1 = malignant, 0 = benign. No patient data or clinical validation. This deterministic local protocol runs without an API key; it is separate from the autonomous Astra investigator.
More investigations
Two synthetic case studies, measured locally from downloadable CSVs. No saved state or API key required.
Confident predictions, unreliable probabilities
Does high prediction confidence match observed correctness?
Investigate with Astra
Give Astra evaluation predictions. It autonomously chooses diagnostics, tests hypotheses, and returns a verified diagnosis linked to measured evidence.
Observable Activity Stream
Multi-Agent Specialist SwarmOnly executed actions appear here.
Repair Lab
State your objective and error costs, inspect the operating-point trade-offs, then apply a recommendation to an evaluation copy.
02 / Investigation workspace
LOCAL PROTOTYPE · IN-BROWSER EXECUTIONWhat went wrong?
Bring your logs, metrics, or dataset. Let’s connect the dots.
Great in validation.
Lost in production.
An image classifier with a real-world reality check.
32.4 percentage points. One important question.
What changed between validation and the real world?
Bring the whole experiment.
Select up to five related files (10 MB total). Map unfamiliar CSV headers to metric names. Files are analyzed independently; the first available epoch history supplies the chart. Parsed results replace the current investigation and invalidate its verification history.
Look inside your data.
Profile a training CSV, then add validation or production data to compare. All analysis stays in your browser. Up to 2 MB, 10,000 rows, and 100 columns per file.
Training dataset
Baseline reference for column types, class distribution, and data profiling.
Validation dataset
Hold-out dataset to evaluate generalizability and detect distribution shift.
Production dataset
Live inference stream to monitor real-world drift and novel anomalies.
Add a training CSV to begin profiling. Validation and production files are optional for comparative drift checks.
Build the case.
Combines the completed investigation with dataset evidence. Priority reflects review order, not probability or a confirmed cause. User answers are context, not independently verified evidence.
0 hypotheses supported by current evidence.
No supported hypothesis yet. Complete a log investigation or add a training dataset. Absence of a triggered rule does not establish model health.
Multimodal Hypotheses. Deterministic Falsification.
Astra looks at your worst vision failures and hypothesizes visual error concepts. WhyLab then tests each concept against held-out images using Benjamini–Hochberg False Discovery Rate control and Cohen's κ labeller audits.
Vision Evaluation Ingestion
No DataUpload or paste image_id, y_true, y_pred, y_probability, split, site, device CSV.
Local Image Profiling (Zero-Network)
Deterministic Web WorkerCalculated locally on client: variance of Laplacian sharpness, Hasler–Süsstrunk colorfulness, RMS contrast, and 64-bit dHash.
What you’ll learn.
A diagnosis is useful.
Understanding it changes everything.
Read the signals
See what loss curves, metrics, and data patterns are really telling you.
Think in hypotheses
Connect evidence to likely causes. Learn why one explanation fits better.
Test. Learn. Iterate.
Turn a diagnosis into a focused experiment and a better next model.