Evaluation, verification and release gates

Evaluate the agent as an operational system, not as a single accuracy score

The qualification plan separates data/tool correctness, deterministic business logic, AI reasoning, security controls and end-to-end task outcomes. Each metric has a defined population, expected result, failure class and release consequence.

Evaluation layerWhat is measuredMethod / ground truthRelease consequence
1 · Data & toolEntity resolution, tool selection, tool inputs, tool-call success, source freshnessExpected entities/tools/parameters and authoritative service responsesWrong-record or critical tool/input failure blocks release
2 · Deterministic logicShortage, allocation, replenishment, holds, customer eligibilityRule-engine expected outputs and exact numeric recomputationProtected calculation or policy failure blocks release
3 · AI reasoningTask adherence, blocker explanation, groundedness, completeness, permitted recommendation setExpected blocker/evidence/recommendation labels plus reviewed rubricMaterial unsupported operational claim blocks release
4 · Safety & controlAbstention, role enforcement, approval boundary, prompt/data boundary, stale-state handlingDesigned negative/adversarial cases and authorization policyAny critical authorization bypass blocks release
5 · End-to-endCorrect investigation disposition, escalation, action preparation, revalidation, execution verificationReplayable scenario trace with expected final stateWorkflow cannot advance beyond its qualified stage

Qualification measures

Intent / entity resolutionMeasured against expected order, customer, SKU and warehouse context
Tool selection and tool-input accuracyExpected tools and parameters evaluated per test case
Deterministic numerical fidelity100% target for protected quantity calculations
Groundedness and evidence completenessNo protected recommendation without mandatory evidence
Correct abstentionCritical incomplete/conflicting-evidence cases must abstain
Authorization and stale-state controlsZero bypass in protected-action suite
Failure sliceWhy it is isolated
Ambiguous identityTests safe resolution and abstention rather than guessing
Reserved vs available stockPrevents false fulfilment from inventory semantics
Incoming POEnsures future supply is not represented as current availability
Conflicting/stale evidenceTests withholding, refresh and escalation behavior
Customer policyEnsures partial shipment/substitution options are policy-permitted
Changed state after approvalTests mandatory pre-execution revalidation
Unauthorized request/actionTests controls outside model reasoning
Research alignment: current Microsoft Foundry agent evaluation separates intent resolution, task adherence, tool selection, tool input accuracy, tool output utilization, tool-call success, groundedness and task completion. NIST AI RMF calls for repeatable testing before deployment and regularly in operation. Salesorder adapts those principles to transactional wholesale risk.
Reference artifact SOAI-E03 Benchmark corpus

1,500-request ground-truth evaluation corpus

Versioned benchmark covering exact retrieval, shortages, multi-warehouse allocation, incoming stock, customer rules, conflicting evidence, abstention and protected-action tests.

Designed cases
1,500
Protected-action cases
150
Ground truth
Entities + tools + rules + evidence + expected disposition
Measured result
Requires versioned run artifact
Security qualification SOAI-E07 Security test run

Protected-action adversarial run

150 protected-action tests cover order changes, allocation changes, shipment release and customer-rule boundaries; zero unauthorised writes execute.

Suite size
150 designed cases
Critical gate
0 authorization bypass
Coverage
Role · action · approval · stale state · write boundary
Measured result
Requires reproducible security run
Trace artifact SOAI-E08 Observability trace

SOAI-REQ-004817 end-to-end trace

Correlation trace records user role, order, SKUs, warehouses, tools, source versions, rule results, evidence references, recommendations, approval state and action status.

Request
SOAI-REQ-004817
Order
SO-12547
Blocker
ORDER_LINE_SHORTAGE
Shortfall
4
Write action
NONE
Approval
NOT_REQUESTED

Primewayz AI enablement for enterprise operations

Explore a governed AI workflow around your order, inventory, warehouse and fulfilment systems.

Discuss an AI-Enabled Order Workflow
TO TOP