01 / Investigation library Local browser storage

Save explicitly to keep parsed evidence, datasets, notes, context answers, and recorded experiments after reload. JSON exports include these data. Browser storage is local to this browser and site.

WhyLabBETA 0.1
AI ML RELIABILITY INVESTIGATOR

Every failed model is trying
to tell you something.

Sentry for software · Datadog for infra · WhyLab for machine learning

Upload your evaluation data. WhyLab investigates why your model is failing, proves the root cause with counterfactual experiments, and repairs your operating policy.

YOUR MODEL SAYS94.2% ACCURACYLooks excellent at first glance
WHYLAB SAYSHIGH-RISK FAILURE DETECTEDMalignant recall: 22.4% · Misses 3 out of 4 cancers
Evidence, not guessworkScientific provenanceFalsification engine
Local deterministic analysisFree, instant, runs completely in your browser without API keys.Astra-powered investigationAutonomous AI diagnostics requiring your own API/deployment token.
  1. 01Observe
    the failure signal
  2. 02Investigate
    with measured evidence
  3. 03Test
    the strongest hypothesis
  4. 04Repair
    and re-evaluate
FLAGSHIP CASE / REPRODUCIBLE LOCAL EXPERIMENT

High accuracy. Missed malignant cases.

A synthetic melanoma classifier exposes the accuracy paradox. Follow the evidence, test competing explanations, then inspect a measured policy repair.

Synthetic teaching data: 1 = malignant, 0 = benign. No patient data or clinical validation. This deterministic local protocol runs without an API key; it is separate from the autonomous Astra investigator.

Validation Accuracy94.2%Looks healthy at a glance
VS
Malignant Recall22.4%Misses 3 out of 4 malignant cases
MEASURED OPERATING-POINT SWEEP

Decision Threshold vs. Malignant Recall & Error Cost

100 evaluation rows · 10 malignant / 90 benign
Malignant Recall and Error Cost across Decision ThresholdsShows that lowering threshold from 0.50 to 0.20 increases malignant recall from 20 percent to 100 percent while dropping error cost from 80 to 4 units.0%25%50%75%100%T=0.0T=0.2T=0.4T=0.6T=0.8T=1.0Baseline (T=0.5)Repaired (T=0.2)
Malignant Recall (%) Total Error Cost (Units)Hover or tap points along curve
OPERATING POINT AT T = 0.20RECOMMENDED
Malignant Recall100.0%10 of 10 detected
Missed Cancers (FN)0 casesZero false negatives
Total Error Cost4 unitsFN×10 + FP×1
Overall Accuracy96.0%96 of 100 correct
Baseline (T=0.50):92.0% Acc · 20.0% Recall · Cost: 80
Repaired (T=0.20):96.0% Acc · 100.0% Recall · Cost: 4
Measured impact: 95% reduction in error cost and +80 pp malignant recall, without sacrificing overall accuracy.
Download source CSV
ADDITIONAL INVESTIGATIONS

More investigations

Two synthetic case studies, measured locally from downloadable CSVs. No saved state or API key required.

Interactive Case

Confident predictions, unreliable probabilities

Does high prediction confidence match observed correctness?

Download case CSV
AUTONOMOUS INVESTIGATION

Investigate with Astra

CHECKING CONNECTION

Give Astra evaluation predictions. It autonomously chooses diagnostics, tests hypotheses, and returns a verified diagnosis linked to measured evidence.

✨ Instant Demo · No Token RequiredExplore a fully worked Astra investigation

See Astra’s multi-step autonomous diagnostic reasoning, hypothesis tests, and verified evidence graph without setting up an API token or incurring OpenAI costs.

Evaluation Evidence & Baseline

Required columns: y_true, y_pred, y_probability. Probability must refer to the positive label. Up to 2 MB per file and 20,000 rows total.

Bounded TXT/LOG/JSON/CSV artifacts up to 16 KB and 200 lines. Astra receives source-linked raw observations plus parsed epoch history.

Quick Directives:

Required for paid Astra requests. Kept only in this page session and sent to the same-origin server.

Observable Activity Stream

Multi-Agent Specialist Swarm

Only executed actions appear here.

    MEASURE → REVIEW → APPLY

    Repair Lab

    State your objective and error costs, inspect the operating-point trade-offs, then apply a recommendation to an evaluation copy.

    Repair Evidence & Cost Matrix

    Binary labels 1/0; probability refers to label 1. Paste up to 2 MB. Baseline threshold must reproduce the supplied predictions.

    Weight for missing true positive cases (e.g. 50x)
    Weight for false alarms (e.g. 1x)
    Target acceptance ceiling
    Current decision boundary (default 0.5)

    Numeric costs are authoritative only after you enter or confirm them; Astra proposals never apply automatically. Total cost = FN × false-negative cost + FP × false-positive cost. Sweep 0–1 in steps of 0.01, retaining the baseline; choose minimum eligible cost, highest threshold on ties.

    Required for live OpenAI API requests. If left blank, verified recorded demonstration values are used.

    02 / Investigation workspace

    LOCAL PROTOTYPE · IN-BROWSER EXECUTION
    START WITH THE EVIDENCE

    What went wrong?

    01

    Bring your logs, metrics, or dataset. Let’s connect the dots.

    Frontend demo · Your evidence stays in your browser
    · CASE FILE / 001SAMPLE INVESTIGATION

    Great in validation.
    Lost in production.

    An image classifier with a real-world reality check.

    Computer visionResNet-1830 epochs
    Validation accuracy94.2%Looking good in the lab
    Production accuracy61.8%A different story outside
    ACCURACY ACROSS ENVIRONMENTS30 epochs
    1007040
    ● Validation● ProductionIllustrative data

    32.4 percentage points. One important question.
    What changed between validation and the real world?

    3 hypotheses. A path to understanding.
    04 / EXTENDED EVIDENCE IMPORT

    Bring the whole experiment.

    Select up to five related files (10 MB total). Map unfamiliar CSV headers to metric names. Files are analyzed independently; the first available epoch history supplies the chart. Parsed results replace the current investigation and invalidate its verification history.

    03 / DATASET INVESTIGATION

    Look inside your data.

    Profile a training CSV, then add validation or production data to compare. All analysis stays in your browser. Up to 2 MB, 10,000 rows, and 100 columns per file.

    Quick-Load Verified Multi-Split Experiment:
    01

    Training dataset

    Empty

    Baseline reference for column types, class distribution, and data profiling.

    02

    Validation dataset

    Empty

    Hold-out dataset to evaluate generalizability and detect distribution shift.

    03

    Production dataset

    Empty

    Live inference stream to monitor real-world drift and novel anomalies.

    No training dataset uploaded

    Add a training CSV to begin profiling. Validation and production files are optional for comparative drift checks.

    05 / COMBINED DIAGNOSIS

    Build the case.

    Combines the completed investigation with dataset evidence. Priority reflects review order, not probability or a confirmed cause. User answers are context, not independently verified evidence.

    0 hypotheses supported by current evidence.

    No supported hypothesis yet. Complete a log investigation or add a training dataset. Absence of a triggered rule does not establish model health.

    MULTIMODAL EXTENSION · COMPUTER VISION INCIDENT INVESTIGATOR

    Multimodal Hypotheses. Deterministic Falsification.

    Astra looks at your worst vision failures and hypothesizes visual error concepts. WhyLab then tests each concept against held-out images using Benjamini–Hochberg False Discovery Rate control and Cohen's κ labeller audits.

    01

    Vision Evaluation Ingestion

    No Data

    Upload or paste image_id, y_true, y_pred, y_probability, split, site, device CSV.

    02

    Local Image Profiling (Zero-Network)

    Deterministic Web Worker

    Calculated locally on client: variance of Laplacian sharpness, Hasler–Süsstrunk colorfulness, RMS contrast, and 64-bit dHash.

    No image profiles loaded. Load the vision flagship case or upload images to profile.
    FAILURE IS A STARTING POINT

    What you’ll learn.

    A diagnosis is useful.
    Understanding it changes everything.

    01

    Read the signals

    See what loss curves, metrics, and data patterns are really telling you.

    02

    Think in hypotheses

    Connect evidence to likely causes. Learn why one explanation fits better.

    03

    Test. Learn. Iterate.

    Turn a diagnosis into a focused experiment and a better next model.