Evaluation & testing

The proof is judged against known industrial records, not fluent prose

The DES-Prime benchmark contains 1,200 versioned requests with expected entity, period, evidence, calculation and action outcomes. Results below are simulated implementation measurements.

Test populationCasesWhat is checked
Exact retrieval & versioning420Correct establishment, period, source and submission version.
Identity / classification ambiguity180Clarification or safe non-merge when evidence is insufficient/conflicting.
Anomaly & correction history220Correct validation rule, previous record, correction chain and state.
Trend / aggregation260Compatible accepted records, exclusions and reproducible calculations.
Protected-action adversarial tests120No unauthorised correction, acceptance or publication action.

Measured simulated benchmark

Evidence-qualified retrieval1,162 / 1,200 (96.8%)
Numerical fidelity1,193 / 1,200 (99.4%)
Ambiguity safely handled177 / 180 (98.3%)
Citation completeness1,176 / 1,200 (98.0%)
Protected actions executed without authority0 / 120
Observed failure classCountRelease response
Wrong evidence/version selected38Added version/status constraints and regression cases.
Numerical mismatch7Blocked response, corrected calculation/mapping path, regression added.
Ambiguity not safely handled3Raised clarification threshold and added adversarial variants.
Citation incomplete24Response cannot pass evidence-complete state until source references attach.
Unauthorised protected action0Zero-tolerance release gate retained.
Because DES-Prime is fictional, these measurements establish reference proof-of-action behaviour only. A real department pilot must rerun the same evaluation design against authorised records and locally approved acceptance thresholds.

Primewayz industrial statistical intelligence

Explore a governed AI workflow around your industrial records, scrutiny process and statistical systems.

Discuss an Industrial Statistics AI Workflow
TO TOP