All certifications / Azure AI Fundamentals / Cheat sheet
Azure AI Fundamentals AI-900 cheat sheet
Domain 1: Describe AI workloads and considerations (19%)
Exam tips
- If a scenario says the outcome is hard to express as rules and there is plenty of historical example data, the answer involves machine learning. Remember that model outputs are probabilistic, so answers claiming AI is always correct are wrong.
- Look for the verb in the scenario: predict or estimate points to ML prediction, flag unusual points to anomaly detection, see or read an image points to computer vision, understand text points to NLP, extract fields from forms points to document processing, and create or draft points to generative AI.
- Count and locate means object detection; one label for the whole image means classification; read text means OCR. If a question mentions identifying who a person is, expect Limited Access to be part of the correct answer.
- Separate the direction of speech tasks: speech to text makes transcripts and captions; text to speech makes audio. Analysis tasks (sentiment, key phrases, entities) never generate new text.
- OCR alone gives you text; document processing gives you fields and tables; knowledge mining gives you a searchable index across many documents. Choose the one that matches the output the scenario wants.
- Copilot: assists a user inside an app. Agent: can also use tools and take actions. Both are generative AI. Watch for the word 'create' or 'draft' as the signal for a generative AI workload.
- Unequal outcomes for groups points to fairness. Testing edge cases, handling failures safely or monitoring for drift points to reliability and safety.
- Data protection, consent, encryption and access control map to privacy and security. Accessibility, languages and designing for people with disabilities map to inclusiveness.
- Explain or disclose points to transparency. Own, govern, review or answer for points to accountability. Both are often paired with Limited Access and transparency notes in questions.
- If a question asks why face identification requires an application, the answer is Limited Access for responsible use, not pricing or preview status. Emotion, gender and age inference are no longer offered by Azure AI Face.
Key terms
- Model
- A function learned from data that takes new input and returns a prediction, label, score or generated content.
- Training
- The process of adjusting a model's internal values using example data until its outputs match the examples well.
- Inference
- Using a trained model to make predictions on new data; also called inferencing or scoring.
- Rule-based software
- Software whose behavior comes from explicit logic written by a developer rather than from patterns learned from data.
- Workload
- A category of problem that AI addresses, such as computer vision, NLP or anomaly detection.
- Anomaly detection
- Identifying data points or events that differ significantly from the normal pattern.
- Knowledge mining
- Extracting information from large amounts of unstructured content and making it searchable.
- Generative AI
- AI that creates new content, such as text, code or images, from a prompt.
- Image classification
- Assigning one or more labels to an image as a whole, without locating objects.
- Object detection
- Finding objects in an image and returning a label and bounding box for each.
- Optical character recognition (OCR)
- Extracting printed or handwritten text from images or scanned documents.
- Bounding box
- The coordinates of a rectangle that marks where an object or word appears in an image.
- Natural language processing (NLP)
- AI that analyzes, understands and generates human language in text form.
- Sentiment analysis
- Determining whether text expresses a positive, negative, neutral or mixed opinion.
- Speech recognition
- Converting spoken audio into text; also called speech to text.
- Conversational AI
- Software such as bots and assistants that interact with users through natural dialog.
- Document intelligence
- AI that extracts text, key-value pairs, tables and named fields from documents into structured data.
- Key-value pair
- A label and its value found in a document, such as 'Due date' and '30 June'.
- Search index
- A store of processed content and extracted fields that users or apps can query quickly.
- Prompt
- The instruction, question or context given to a generative AI model to produce a response.
- Large language model (LLM)
- A very large neural network trained on text that generates language by predicting tokens.
- Copilot
- A generative AI assistant embedded in an app that helps a user with tasks while the user stays in control.
- Agent
- A generative AI solution that combines a model with instructions, knowledge and tools so it can take actions toward a goal.
- Fairness
- The principle that AI should treat all people fairly and not give different outcomes to similar people based on irrelevant characteristics.
- Reliability and safety
- The principle that AI should perform consistently and safely as intended, including in unexpected conditions.
- Bias
- A systematic skew in data or model behavior that leads to unfair outcomes for some groups.
- Model drift
- A decline in a model's performance over time as real-world data changes from the training data.
- Privacy and security
- The principle that AI systems should protect personal data and be secure against misuse and attack.
- Inclusiveness
- The principle that AI should empower everyone and be designed to be usable by people of all abilities and backgrounds.
- Personally identifiable information (PII)
- Data that can identify a person, such as a name, phone number, address or ID number.
- Data minimization
- Collecting and keeping only the personal data that a purpose actually needs.
- Transparency
- The principle that people should understand how an AI system works, what it is for and what its limitations are.
- Accountability
- The principle that the people who design and deploy AI systems are answerable for how they operate.
- Explainability
- Showing which inputs or features most influenced a model's output.
- Transparency note
- Microsoft documentation describing an AI service's capabilities, intended uses and limitations.
- Impact assessment
- A documented review of an AI system's intended uses, stakeholders, potential harms and mitigations.
- Human in the loop
- A design where a person reviews or approves AI outputs before they take effect.
- Limited Access
- Microsoft's policy that requires approval of the customer and use case before sensitive AI features can be used.
- Custom neural voice
- A Speech capability that creates a synthetic voice resembling a specific person; it is a Limited Access feature.
Domain 2: Fundamental principles of machine learning on Azure (19%)
Exam tips
- If the question asks which column is the label, pick the value being predicted. A big gap between training and validation performance means overfitting.
- Numeric label means regression. Metrics with 'error' in the name should be low; R² should be close to 1. Precision and recall are never regression metrics.
- Two classes is binary, more is multiclass, even if the classes are numbers like ratings 1 to 5. Missing positives is costly: recall. False alarms are costly: precision.
- No labels and 'discover groups' means clustering. Known categories to predict means classification, even if the scenario uses the word 'group'.
- Deep learning is a subset of machine learning, which is a subset of AI. If a scenario involves images, audio or natural language at scale, deep learning is the technique underneath.
- Tokens are the units of text; embeddings are the vectors that represent meaning; attention relates tokens to each other in context. GPT-style models generate by predicting the next token.
- Compute instance is for development by one person; compute cluster is for scalable training jobs. Azure Machine Learning is for training your own models; Azure AI services give you prebuilt ones.
- Try many algorithms automatically and pick the best: AutoML. Build a pipeline visually by connecting components: designer. Neither requires writing code.
- Someone waiting for an immediate answer means a real-time (online) endpoint. Scoring a large dataset on a schedule means batch. Questions about finding where a model makes more errors or treats groups differently point to the Responsible AI dashboard.
Key terms
- Feature
- An input value the model uses to make a prediction, such as age or floor area.
- Label
- The value a supervised model is trained to predict, such as price or a yes/no outcome.
- Validation data
- Data held back from training and used to measure how well the model performs on unseen examples.
- Overfitting
- When a model learns the training data too closely and performs poorly on new data.
- Regression
- Supervised learning that predicts a numeric value.
- Mean absolute error (MAE)
- The average absolute difference between predicted and actual values, in the label's units.
- Root mean squared error (RMSE)
- The square root of the average squared error; it penalizes large errors more than MAE.
- Coefficient of determination (R²)
- The proportion of variance in the label explained by the model; closer to 1 is better.
- Binary classification
- Predicting one of two classes, such as yes or no.
- Confusion matrix
- A table of actual versus predicted classes showing true and false positives and negatives.
- Precision
- Of the items predicted positive, the proportion that were actually positive: TP / (TP + FP).
- Recall
- Of the actual positive items, the proportion the model found: TP / (TP + FN).
- Clustering
- Unsupervised learning that groups items with similar feature values.
- Unsupervised learning
- Machine learning on data without labels, which finds structure on its own.
- Supervised learning
- Machine learning on data that includes known labels, used to predict those labels.
- k-means
- A clustering algorithm that groups items around k center points, repeatedly adjusting the centers.
- Neural network
- A model made of layers of connected artificial neurons whose weights are learned during training.
- Weight
- A number on a connection between neurons that is adjusted during training to reduce error.
- Deep learning
- Machine learning with neural networks that have many hidden layers.
- Loss function
- A calculation that measures how far a model's predictions are from the correct answers.
- Token
- A unit of text, such as a word or part of a word, that a language model processes.
- Embedding
- A vector of numbers representing the meaning of a token or text, where similar meanings are close together.
- Attention
- A mechanism that lets each token weigh the relevance of the other tokens in the sequence.
- Transformer
- A neural network architecture built on attention, used by modern language models.
- Workspace
- The top-level Azure Machine Learning resource that holds data, compute, jobs, models and endpoints.
- Azure Machine Learning studio
- The web portal for working with an Azure Machine Learning workspace.
- Compute instance
- A managed development VM for one user's notebooks and experiments.
- Compute cluster
- A group of VMs that scales automatically for training jobs and can scale to zero when idle.
- Automated machine learning (AutoML)
- A feature that tries many algorithms and settings automatically and ranks the resulting models by a chosen metric.
- Primary metric
- The measure AutoML uses to rank models, such as accuracy or normalized RMSE.
- Designer
- A drag-and-drop canvas in Azure Machine Learning studio for building training pipelines visually.
- Featurization
- Preparing raw data for training, such as handling missing values and encoding categories.
- Endpoint
- A web address where a deployed model accepts input data and returns predictions.
- Online (real-time) endpoint
- An endpoint that returns predictions immediately for individual requests.
- Batch endpoint
- An endpoint that scores large volumes of data asynchronously as a job and writes results to storage.
- Responsible AI dashboard
- An Azure Machine Learning tool combining error analysis, fairness, interpretability and what-if analysis for a model.
Domain 3: Computer vision workloads on Azure (19%)
Exam tips
- Whole image, one answer: classification. Where and how many: object detection with bounding boxes. Exact outline, pixel by pixel: segmentation.
- CNNs learn their filters during training rather than using hand-chosen ones. A model that can answer questions about images or match images to text is multimodal.
- Sentence describing the image: captions. List of keywords: tags. Where things are: object or people detection. Thumbnail region: smart crops. Text in the image: OCR (Read).
- Extract text from images: OCR with Read. Extract named fields or tables from forms: Document Intelligence. Describe the image in words: captions, not OCR.
- Detection and image attributes such as blur, glasses and head pose: available. Identification and verification: Limited Access. Emotion, gender and age: retired.
- Common business document with standard fields: prebuilt model. Tables and structure from any document: layout. Your own unique form: custom model trained on labeled samples.
- One key and endpoint for many services, one bill: multi-service. Free F0 tier or separate billing for one service: single-service. An app needs the endpoint plus a key (or Entra ID) to call it.
- Map output to service: description or objects to Vision image analysis, text to Vision Read, fields and tables to Document Intelligence, faces to Face, your own classes to a custom model.
Key terms
- Image classification
- Predicting one or more labels for an image as a whole.
- Object detection
- Locating each object in an image with a class label and bounding box.
- Semantic segmentation
- Classifying every pixel in an image to produce a precise mask of each class.
- Confidence score
- A value between 0 and 1 showing how sure the model is about a prediction.
- Pixel
- The smallest element of a digital image, stored as one or more numeric values.
- Filter (kernel)
- A small grid of weights applied across an image to produce a feature map, for example to highlight edges.
- Convolutional neural network (CNN)
- A deep learning model that learns filters to extract features from images for tasks such as classification.
- Multimodal model
- A model trained on more than one type of data, such as images and text together.
- Caption
- A generated sentence describing an image, returned with a confidence score.
- Dense captions
- Captions for multiple regions of an image, each with a bounding box.
- Tag
- A word describing something visible in an image, such as an object, setting or action.
- Smart crop
- A suggested crop region that keeps the area of interest for a given aspect ratio.
- Optical character recognition (OCR)
- Extracting printed or handwritten text from images and documents.
- Read feature
- The Azure AI Vision OCR capability that returns lines and words with their positions and confidence.
- Bounding polygon
- The set of coordinates outlining where a line or word appears in the image.
- Handwriting recognition
- OCR of handwritten rather than printed text.
- Face detection
- Finding faces in an image and returning their location and image attributes.
- Face verification
- Checking whether two face images belong to the same person (one-to-one).
- Face identification
- Finding which known person a face belongs to from a group (one-to-many).
- Facial landmarks
- Points on a face, such as eye corners and nose tip, returned by face detection.
- Prebuilt model
- A Document Intelligence model trained by Microsoft for a common document type, such as invoices or receipts.
- Layout model
- A model that extracts text, tables, selection marks and structure from any document.
- Custom extraction model
- A model you train with labeled samples to extract your own fields from your own document type.
- Selection mark
- A checkbox or radio button on a form, returned as selected or unselected.
- Multi-service resource
- An Azure AI services resource that gives one endpoint and set of keys for several AI services with one bill.
- Single-service resource
- A resource for one AI service, such as Azure AI Vision, often with a free F0 tier.
- Endpoint
- The URL an application calls to use the AI service.
- Resource key
- A secret value sent with requests to authenticate to an AI service; each resource has two.
- Azure AI Vision
- The Azure AI service for image analysis and OCR with prebuilt models.
- Azure AI Face
- The Azure AI service for detecting and analyzing faces, with recognition under Limited Access.
- Azure AI Document Intelligence
- The Azure AI service that extracts fields, tables and structure from documents.
- Custom vision model
- An image model trained on your own labeled images to recognize classes that prebuilt models do not cover.
Domain 4: Natural language processing workloads on Azure (19%)
Exam tips
- Overall mood of text: sentiment analysis. Mood toward specific features mentioned: opinion mining. Which language: language detection. None of them translate.
- Main topics: key phrases. Categorized mentions such as people and dates: NER. Disambiguate and link to Wikipedia: entity linking. Find and mask personal data: PII detection.
- Counting words cannot tell that 'car' and 'automobile' are related; embeddings can. Semantic search and similarity questions point to embeddings.
- Short version of a long text: summarization (extractive picks sentences, abstractive writes new ones). Your own categories: custom text classification trained on labeled examples.
- Always check the direction: captions and transcripts are speech to text; reading text aloud is text to speech. Custom neural voice needs Limited Access approval.
- Text or documents: Azure AI Translator. Spoken audio: speech translation in Azure AI Speech. Converting scripts without changing meaning: transliteration.
- In 'Order two large pizzas', the whole sentence is the utterance, OrderPizza is the intent, and 'two' and 'large' are entities. CLU returns structured intent and entities, not an answer.
- Existing FAQ content and 'answer common questions' points to question answering. 'Understand what the user wants to do and extract details' points to conversational language understanding.
- Text analysis: Language. Audio: Speech. Different language: Translator, unless the input is spoken, then speech translation. Action from a command: CLU. Answer from FAQs: question answering.
Key terms
- Azure AI Language
- The Azure AI service that analyzes and understands text with prebuilt and customizable features.
- Language detection
- Identifying the language of text and returning its name, ISO code and a confidence score.
- Sentiment analysis
- Labeling text or sentences as positive, neutral, negative or mixed with confidence scores.
- Opinion mining
- Aspect-based sentiment that links opinions to specific targets mentioned in the text.
- Key phrase extraction
- Returning the main talking points of a text as a list of phrases.
- Named entity recognition (NER)
- Finding entities in text and classifying them as types such as Person, Location or Organization.
- Entity linking
- Identifying which known real-world entity a mention refers to and linking it to a knowledge base entry.
- PII detection
- Finding personal information in text and returning a redacted version.
- Tokenization
- Splitting text into tokens such as words or subword pieces for a model to process.
- TF-IDF
- A weighting that scores words by how frequent they are in a document and how rare across the collection.
- Embedding
- A vector that represents the meaning of text so similar meanings are close together.
- Semantic similarity
- How close two texts are in meaning, often measured with cosine similarity between embeddings.
- Extractive summarization
- Summarizing by selecting and returning the most important sentences from the original text.
- Abstractive summarization
- Summarizing by generating new sentences that capture the main ideas.
- Custom text classification
- Training Azure AI Language to assign your own categories to documents using labeled examples.
- Multi-label classification
- Classification in which one document can receive more than one label.
- Speech to text
- Converting spoken audio into written text; also called speech recognition.
- Text to speech
- Converting text into spoken audio; also called speech synthesis.
- Neural voice
- A natural-sounding synthetic voice produced by a deep learning model.
- SSML
- Speech Synthesis Markup Language, used to control pronunciation, rate, pitch and style of synthesized speech.
- Neural machine translation
- Translation by deep learning models that consider the whole sentence and its context.
- Transliteration
- Converting text from one writing script to another without translating its meaning.
- Document translation
- Translating whole files while preserving their structure and formatting.
- Speech translation
- Translating spoken audio into text or speech in another language in near real time.
- Utterance
- An example of something a user might say or type to an app.
- Intent
- The goal or action a user wants, predicted from an utterance.
- Entity
- A specific detail in an utterance that the app needs, such as a date, place or quantity.
- Conversational language understanding (CLU)
- An Azure AI Language feature that predicts intents and extracts entities from user input.
- Question answering
- An Azure AI Language feature that answers natural language questions from a knowledge base of question and answer pairs.
- Knowledge base
- The collection of question and answer pairs, often imported from FAQs and documents, that question answering searches.
- Chit-chat
- Prebuilt responses to small talk that give a bot a consistent personality.
- Multi-turn conversation
- Follow-up prompts that guide a user through several steps within a topic.
- Azure AI Speech
- Service for speech to text, text to speech and speech translation.
- Azure AI Translator
- Service for translating text and documents between languages.
- Service chaining
- Combining several AI services in sequence so one's output becomes the next one's input.
Domain 5: Generative AI workloads on Azure (24%)
Exam tips
- LLMs predict tokens; they do not look up stored answers. Anything about knowledge limits, cutoff dates or fabricated facts links back to this, and the usual fix is grounding with your own data.
- Create, draft, rewrite or converse freely: generative AI. Extract a specific field or assign a fixed label reliably: a traditional AI service may be the better answer.
- Role, rules and format that apply to the whole conversation go in the system message. Examples in the prompt are few-shot. Changing the model's weights is fine-tuning, not prompt engineering.
- Model must answer from your current or private data without retraining: RAG. Model must adopt a style or format consistently: fine-tuning may help. RAG retrieves first, then generates.
- Need repeatable, factual output: lower temperature. Need creative variety: raise it. Answers cut off mid-sentence: max tokens is too low. Adjust temperature or top_p, not both.
- Compare and choose models: model catalog. Try prompts and settings without code: playground. Organize a solution's assets and team access: project. Expect both the Azure AI Foundry and Microsoft Foundry names.
- Vectors for search and RAG: embeddings model. Conversation and text generation: GPT chat model. Pictures from text: image generation model. Azure OpenAI is billed, with no free tier.
- Only answers questions: chat assistant or copilot. Uses tools to take actions or complete multi-step tasks: agent. Agents need least-privilege tools and human approval for high-impact actions.
- Memorize the order: identify, measure, mitigate, operate. Mitigation has four layers: model, safety system, system message and grounding, user experience. Content filters cover hate, sexual, violence and self-harm.
- Supported by sources: groundedness. Answers the question: relevance. Reads naturally: fluency. Logically organized: coherence. Deliberately trying to break it: red teaming.
Key terms
- Large language model (LLM)
- A very large transformer model trained on massive text data that generates language by predicting tokens.
- Next-token prediction
- Generating text by repeatedly predicting a likely next token and appending it.
- Pretraining
- Initial self-supervised training on large text collections that teaches a model language patterns and knowledge.
- Context window
- The maximum number of tokens a model can handle in one request, covering prompt and response.
- Chat assistant
- A generative AI app that converses with users in natural language over multiple turns.
- Copilot
- A generative AI assistant embedded in an app to help users with tasks while they stay in control.
- Summarization
- Condensing long content into a shorter version that keeps the key points.
- Code generation
- Using a model to write, explain or convert code from natural language descriptions.
- System message
- Instructions sent before the conversation that set the model's role, rules, tone and output format.
- User prompt
- The user's request or question sent to the model.
- Few-shot prompting
- Including a few examples of input and desired output in the prompt so the model follows the pattern.
- Zero-shot prompting
- Giving the model only an instruction, with no examples.
- Grounding
- Providing relevant trusted information in the prompt so the model bases its answer on it.
- Retrieval augmented generation (RAG)
- A pattern that retrieves relevant content from your data and adds it to the prompt before the model generates an answer.
- Vector index
- A search index that stores embeddings so content can be retrieved by semantic similarity.
- Fine-tuning
- Further training a pretrained model on your own examples to change its behavior or style.
- Temperature
- A setting that controls randomness in token selection; low is focused and consistent, high is varied and creative.
- Top_p
- A setting that limits token choices to the most probable set whose combined probability reaches p.
- Max tokens
- A limit on the number of tokens the model can generate in a response.
- Stop sequence
- Text that tells the model to stop generating when it is produced.
- Microsoft Foundry
- Microsoft's platform and portal for building generative AI apps and agents, formerly Azure AI Studio and Azure AI Foundry.
- Project
- A Foundry workspace that holds a solution's model deployments, agents, data connections and evaluations.
- Model catalog
- The Foundry library for discovering, comparing and deploying models from Microsoft, OpenAI and other providers.
- Model card
- Documentation describing a model's capabilities, intended uses, limitations and deployment options.
- Azure OpenAI
- OpenAI models hosted in Azure with Azure security, networking, content filtering and data protection.
- Chat completion model
- A model that takes a conversation of messages and generates the next response, such as a GPT model.
- Embeddings model
- A model that converts text into vectors for semantic search and similarity, not readable text.
- Deployment
- An instance of a model in your resource, with a name and endpoint that your app calls.
- AI agent
- A generative AI application that uses a model with instructions, knowledge and tools to reason and take actions toward a goal.
- Tool
- A capability an agent can call, such as search, code execution or an API.
- Foundry Agent Service
- The Microsoft Foundry capability for building, deploying and managing agents.
- Prompt injection
- Malicious instructions hidden in input content that try to make a model or agent act against its instructions.
- Identify, measure, mitigate, operate
- Microsoft's four stages for developing and running generative AI responsibly.
- Content filter
- A safety system that classifies prompts and responses for harmful content and blocks it above a set severity.
- Azure AI Content Safety
- A service that detects harmful content in text and images and offers features such as prompt attack detection.
- Jailbreak
- A prompt designed to trick a model into ignoring its instructions or safety rules.
- Groundedness
- How well a response's claims are supported by the provided source context.
- Relevance
- How well a response addresses the user's question.
- AI-assisted evaluation
- Using a model as a judge to score responses against criteria such as groundedness or coherence.
- Red teaming
- Deliberately probing an AI system to find harmful outputs and weaknesses before attackers or users do.
Study Azure AI Fundamentals for free
Lessons, quizzes, exam simulations and hands-on labs.
Open the Azure AI Fundamentals study planLessons, quizzes, exam simulations and hands-on labs.