Traditional ML models are evaluated with metrics like accuracy or RMSE against known labels. Generative AI output is open-ended text, so there is rarely one right answer to compare with. Evaluation therefore uses a mix of quality metrics, safety metrics, human review and adversarial testing, and it continues after release.
Keep reading for free
Create a free StudyToCert account to read the rest of this lesson: 5 more sections, 4 key terms, a real-world example, an exam tip and self-check questions. Every lesson, lab and practice test is free with an account.