Qualification and outcomes
Prove technical control first, then measure whether the workflow improves wholesale operations
Reference evaluation and production business outcomes are deliberately separated. A synthetic run can qualify correctness and control; it cannot establish live productivity, service-level or ROI improvement.
Reference qualification model
Salesorder AI qualification and production measurement framework
The reference implementation separates reproducible technical qualification from live business outcomes. Synthetic benchmark values are publishable only when backed by a versioned run artifact; productivity and ROI remain production measurements.
Operating assumptions
- The AI enablement environment uses synthetic transactional records modelled on wholesale ERP workflows.
- The benchmark contains 1,500 versioned requests with known expected evidence, calculations, permissions and outcomes.
- Salesorder transactional services remain the source of truth; the model does not become an order, inventory or warehouse system of record.
- Production productivity, service-level and ROI outcomes require validation with authorised live data and users.
Qualification matrix
| Measure | Method | Population | Release gate | Production use |
|---|---|---|---|---|
| Entity / context resolution | Exact expected order, customer, SKU and warehouse scope | Applicable ground-truth cases | Threshold defined before run; ambiguity must abstain | Detect retrieval drift and wrong-record risk |
| Tool selection & input accuracy | Expected tool set and validated parameters per test case | Tool-bearing evaluation cases | Critical tools and parameters must pass | Detect orchestration regressions |
| Numerical fidelity | Deterministic recomputation of outstanding, available, allocated and shortage quantities | Numeric cases | 100% target for protected calculations | Detect quantity/calculation defects |
| Blocker classification | Expected exception class vs system classification | Known exception cases | Predefined acceptance threshold; critical classes separately gated | Monitor operational diagnosis quality |
| Evidence completeness | Mandatory evidence references present and source versions captured | Evidence-bearing cases | 100% target for protected recommendations | Detect unsupported explanations |
| Grounded explanation | Claims trace to retrieved evidence and deterministic rule results | Explanation cases | No material unsupported operational claim | Monitor hallucination / grounding risk |
| Correct abstention | System withholds conclusion when identity, evidence or authority is incomplete/conflicting | Designed abstention cases | Critical abstention cases must pass | Monitor unsafe overreach |
| Authorization enforcement | Protected requests and actions checked against authenticated role and action policy | 150 protected-action tests plus production protected actions | 0 authorization bypasses | Security control effectiveness |
| Pre-execution revalidation | Order/inventory/rule versions re-read before approved write | All prepared protected actions | 100% required | Prevent stale-state execution |
| Trace completeness | Required identity, source, tool, rule, recommendation, approval and action fields captured | All evaluation and production runs | 100% for protected workflows | Auditability and incident reconstruction |
Security and governance qualification targets
Protected action without authenticated authority0 permitted
Approval token accepted after relevant state/version changes0 permitted
Direct model write to order, inventory or shipment tables0 permitted
Protected execution without correlation/audit trace0 permitted
Agent evaluation qualification targets
Intent / entity resolutionMeasured against expected order, customer, SKU and warehouse context
Tool selection and tool-input accuracyExpected tools and parameters evaluated per test case
Deterministic numerical fidelity100% target for protected quantity calculations
Groundedness and evidence completenessNo protected recommendation without mandatory evidence
Correct abstentionCritical incomplete/conflicting-evidence cases must abstain
Authorization and stale-state controlsZero bypass in protected-action suite
Act with Approval qualification targets
Approved actions revalidated immediately before execution100%
Direct model writes to transactional tables0
Bounded automation qualification targets
Read-only shadow review completed before action permission expansionRequired gate
Material exception silently resolved without policy/authority0
Learn & Improve qualification targets
Confirmed defects converted to versioned regression cases100%
Critical authorization regression allowed to release0
| Business KPI | Pre-AI baseline | Controlled rollout measure | How change is calculated |
|---|---|---|---|
| Median exception investigation time | Measured from representative human cases | Measured from assisted cases with same start/stop definition | Pilot median vs baseline median |
| Records/screens manually consulted | Observed/telemetry count per case | Manual views remaining after evidence assembly | Mean/median reduction by exception class |
| First-pass blocker accuracy | Reviewed human disposition sample | Reviewed assisted investigation sample | Correct first disposition / reviewed cases |
| Escalation rate | Escalated exceptions / investigated exceptions | Same denominator during controlled rollout | Rate change with severity slice |
| Rework / reopen rate | Reopened or corrected resolved exceptions / resolved cases | Same definition in pilot | Rate change; lower is better |
| Exception ageing | Time from exception open to permitted next action | Same definition in pilot | Median and tail percentile change |
| Unauthorized protected actions | Production audit count | Production audit count | Must remain zero |
| Evidence-complete investigations | Cases containing required source/version references | Same definition in pilot | Completeness rate |
| Rollout stage | What is allowed | Evidence required to progress |
|---|---|---|
| 0 · Baseline | Human process only; instrument representative exceptions | Stable KPI definitions and reviewed ground truth |
| 1 · Read-only shadow | Agent investigates; no operational recommendation/action dependency | Agreement with reviewed outcomes; failure analysis complete |
| 2 · Assisted investigation | Users receive evidence and explanation | Quality/security gates pass; user review captured |
| 3 · Recommendation | Policy-permitted options shown | Unsupported recommendation and abstention gates pass |
| 4 · Prepare action | Action package created but not executed | Approval path, scope and idempotency validated |
| 5 · Controlled execution | Approved bounded action after live-state revalidation | Security suite, stale-state tests, audit and rollback/recovery controls pass |
Production outcome claims require a defined population, period, baseline, measurement method and source. Until those exist, the case reports qualification status rather than invented productivity percentages.
Primewayz AI enablement for enterprise operations