Functional tests can prove that an AI application responds. They do not necessarily prove that the agent understood the customer, followed the business rule, selected the correct action or recovered safely when something went wrong.
Start with the business journey
Define the outcome before defining the test. A journey such as changing a delivery address, reporting a sensitive event or requesting account help should have explicit expected behaviour, acceptable alternatives and escalation rules.
Test the complete execution
Follow the request through intent or classification, retrieval, tool selection, parameters, orchestration, response generation and handover. The final message is only one observable part of the system.
Evaluate more than exact text
Use criteria for correctness, completeness, relevance, tone, safety and business-rule compliance. For probabilistic outputs, evaluate meaning and evidence rather than requiring one exact sentence.
Turn failures into regression
Every important failure should become a repeatable scenario. The goal is not a one-time report; it is a growing safety net that protects future prompt, model, flow and integration changes.
Need this applied to your AI system?
Turn the principle into an assessment, test strategy or engineering engagement.