A/B Test and Wild Card C
Compare two versions of an OMNIA output on the same inputs, judged blind, and let a third arm written by OMNIA try to beat both. Every judgment is kept until you delete it.
Turn your data into a tested improvement.
Upload CSV or JSON, freeze Wild Card C, compare unseen references, and publish a traceable dataset. Check fresh data later to see whether the improvement holds.
Owner sign-in
Experiments hold your prompts and settings, so this page is for the owner only. Enter the owner key once; this browser remembers it.
Two variants, one input, judged blind
Each variant changes settings only: a prompt, a model, a weight profile. A winner is named only when an always-valid test supports it, so refreshing this page cannot create one.
| Arms | Judged by | Pairs | Wins | p | Needs p ≤ |
|---|
Judge the next pair
No pairs are waiting for you.
Run a trial
Arms and settings
New experiment
Evaluation Lab
Freeze A/B/C, compare on fresh unseen cases, and publish traceable outputs only after they pass. Local numeric fusion, no model API credits.
Open Evaluation Lab →Wild Card C
A third arm that OMNIA writes itself. It blends the settings that are winning, weighs every judgment by how believable its evidence is, and makes one deliberate change each round.
Loading…
Chance of being best
Lineage
No C has been written yet.
Settings to run for
This experiment's outputs come from outside the engine. Run these settings and include C's output under its name when you add a trial.
How the algorithm works
- Evidence. Each judgment counts by the OMNIA evidence hierarchy: a measured outcome , your blind pick , a model judge that agreed with itself in both orders . A judge that contradicts itself counts for nothing.
- Blend. C starts from every arm so far, each weighted by its chance of being best. Numbers and weight profiles are averaged. Text, choices and lists come from the leader.
- One change. C then moves the setting that most separates the leader from the laggard further toward the leader, in steps that shrink as evidence grows, and prefers settings no earlier C has moved. With nothing numeric to move, it toggles one list item, or has the model rewrite the prompt with one stated change.
- Fair traffic. Top-two Thompson sampling picks the next pair to judge, and at least of trials still compare A with B directly.
- Retirement. After weighted judgments, a C with less than a chance of being best is retired, and the next C is written from all the evidence.
- Stricter verdict. Each C version adds two comparisons, and the significance threshold is divided by their number.
C changes settings only. The scoring code is the same for every arm.
Claim IntelligenceTrace the statement. Inspect the sources. Follow the outcome.
See what connects.
One continuous path from a public statement to reusable knowledge.

Give your data a track record.
Track a claim to see real sources and outcomes connected here. No example findings are mixed into your data.
Owner-reviewed outcomes · Connections show recorded lineage, not causation. Motion is decorative; it does not represent measured EFVT or live transfers.