Generative AI output is not simply right or wrong, and small prompt changes can improve one kind of question while breaking another. Evaluation means running your app against a set of test inputs and scoring the outputs systematically, so decisions about prompts, models and retrieval are based on evidence rather than a few chats in the playground. Azure AI Foundry provides built-in evaluators and an evaluation dashboard, and the Azure AI Evaluation SDK lets you run the same evaluators in code or in a pipeline.
Keep reading for free
Create a free StudyToCert account to read the rest of this lesson: 5 more sections, 4 key terms, a real-world example, an exam tip and self-check questions. Every lesson, lab and practice test is free with an account.