Evaluate the agent as an operational system, not as a single accuracy score
The qualification plan separates data/tool correctness, deterministic business logic, AI reasoning, security controls and end-to-end task outcomes. Each metric has a defined population, expected result, failure class and release consequence.
Replayable scenario trace with expected final state
Workflow cannot advance beyond its qualified stage
Qualification measures
Intent / entity resolutionMeasured against expected order, customer, SKU and warehouse context
Tool selection and tool-input accuracyExpected tools and parameters evaluated per test case
Deterministic numerical fidelity100% target for protected quantity calculations
Groundedness and evidence completenessNo protected recommendation without mandatory evidence
Correct abstentionCritical incomplete/conflicting-evidence cases must abstain
Authorization and stale-state controlsZero bypass in protected-action suite
Failure slice
Why it is isolated
Ambiguous identity
Tests safe resolution and abstention rather than guessing
Reserved vs available stock
Prevents false fulfilment from inventory semantics
Incoming PO
Ensures future supply is not represented as current availability
Conflicting/stale evidence
Tests withholding, refresh and escalation behavior
Customer policy
Ensures partial shipment/substitution options are policy-permitted
Changed state after approval
Tests mandatory pre-execution revalidation
Unauthorized request/action
Tests controls outside model reasoning
Research alignment: current Microsoft Foundry agent evaluation separates intent resolution, task adherence, tool selection, tool input accuracy, tool output utilization, tool-call success, groundedness and task completion. NIST AI RMF calls for repeatable testing before deployment and regularly in operation. Salesorder adapts those principles to transactional wholesale risk.