

Make better data.
Prove the difference.
Compare equal weighting, your believability weights, and a frozen Wild Card C. Freeze the methods first. Open fresh test data once. Publish only what passes. Recheck with fresh observations as conditions change.
Local execution · no model API calls · owner access required
1. Define what better means
For numeric estimates in a shared unit: sensor fusion, model predictions, forecasts or extracted media features. This adapter cannot directly grade prose or accept raw audio. Use independently observed reference targets.
2. Open fresh test data
All three approaches will see exactly the same cases. This test opens once. Previously used inputs and groups are rejected across this lab.
3. The same test. Three approaches.
Lower reference error is better. Processing time covers local prediction only. API spend is $0; hosting and CPU are not priced. Changed outputs are not automatically new knowledge.
4. Publish with proof attached
Paired comparisons + corrected evidence thresholds
EFVT: measured descriptors and their limits
Frozen methods + lineage
Inspect case-level results (first 100)
| Case | Reference | A | B | C |
|---|
5. Keep the proof current
Upload a later batch of independently measured references. The frozen methods stay unchanged. Fresh failures flag the published method for review; they never rewrite the original test.
At least 30 fresh independent cases per batch, later than the preceding window. Up to 20 checks per frozen version. Each check uses a stricter evidence threshold. Connectors can submit the same checks through the owner-authorized API.