StudyToCert

All certifications / Azure AI Engineer / Cheat sheet

Azure AI Engineer AI-102 cheat sheet

Every exam tip and key term from the free Azure AI Engineer lessons, by domain. Use your browser's Print to save it as a PDF.

Domain 1: Plan and manage an Azure AI solution (24%)

Exam tips

Key terms

Azure AI services
Microsoft's family of prebuilt AI APIs for vision, language, speech, translation, documents and content safety, called over REST or SDKs.
Azure OpenAI
Azure-hosted OpenAI models for chat, reasoning, embeddings and image generation, deployed and managed through Azure AI Foundry.
Document Intelligence
A service that extracts text, structure and named fields from documents using prebuilt or custom models.
Content Safety
A service that detects harmful text and images across hate, sexual, violence and self-harm categories.
Single-service resource
An Azure resource for one AI service, with its own endpoint, keys, pricing tier and bill.
Multi-service resource
An Azure AI services resource that exposes many AI services behind one endpoint and key pair with combined billing.
Hub
A shared Azure AI Foundry container, built on Azure Machine Learning, that holds connections, security and compute for several projects.
Project
A Foundry workspace where you deploy models, build agents, evaluate and manage the assets for one app or team.
Deployment
A named instance of a model version inside an Azure OpenAI or Foundry resource, called by name from code.
Tokens per minute (TPM)
The rate limit and quota unit for pay-as-you-go deployments.
Provisioned throughput unit (PTU)
A unit of reserved model processing capacity for provisioned deployments.
Data zone
A group of regions, such as the EU or US, within which Data Zone deployments process requests.
kind
The template property that sets which AI service a Microsoft.CognitiveServices/accounts resource provides.
Bicep
A domain-specific language for Azure infrastructure as code that compiles to ARM JSON templates.
customSubDomainName
The property that gives a resource a unique endpoint name, required for Entra ID authentication and private endpoints.
disableLocalAuth
A property that turns off key-based authentication so only Microsoft Entra ID tokens work.
Endpoint
The base URL of an Azure AI resource that all API calls are sent to.
Ocp-Apim-Subscription-Key
The HTTP header that carries a resource key for most Azure AI services.
DefaultAzureCredential
An Azure Identity class that finds a Microsoft Entra ID credential automatically from the environment, managed identity or developer sign-in.
Operation-Location
The header returned by asynchronous operations that points to the URL you poll for the result.
Managed identity
An Entra ID identity for an Azure resource whose credentials Azure creates and rotates automatically.
Cognitive Services User
An RBAC role that lets an identity call Azure AI services data-plane APIs.
Cognitive Services OpenAI User
An RBAC role that allows inference calls to Azure OpenAI deployments without management rights.
Custom subdomain
A unique resource-specific endpoint name required for Entra ID token authentication.
Azure Key Vault
A managed service for storing and auditing secrets, keys and certificates.
Key rotation
Replacing keys regularly, using the two-key pattern so clients never lose access.
Private endpoint
A network interface with a private IP in your virtual network that connects privately to an Azure resource.
Customer-managed key
An encryption key you own in Key Vault used to encrypt a service's data at rest instead of a Microsoft-managed key.
Connected container
A container that processes data locally but reports usage to Azure for billing.
Disconnected container
A container approved for fully offline use under a commitment plan, with no online usage reporting.
Billing setting
The container startup value that holds the endpoint of the Azure resource charged for usage.
Eula=accept
The required startup argument confirming acceptance of the container's license terms.
Diagnostic setting
A resource configuration that routes logs and metrics to Log Analytics, storage or an event hub.
Action group
A reusable set of notifications and automated actions triggered by alerts.
Commitment tier
A pricing plan with a fixed monthly fee for a set volume of usage at a discounted rate.
Budget
A Cost Management threshold on actual or forecasted spend that sends alerts.
Harm categories
The four Content Safety classifications: hate, sexual, violence and self-harm.
Severity level
A score showing how harmful content is within a category, which your app compares to a threshold.
Blocklist
A custom list of terms that Content Safety flags in addition to its harm classifiers.
Transparency
The responsible AI principle that people should know they are using AI and understand its capabilities and limits.
Content filter
The configurable input and output moderation applied to every Azure OpenAI deployment.
Prompt shields
A Content Safety feature that detects user prompt attacks (jailbreaks) and document attacks (indirect prompt injection).
Groundedness detection
A check that flags model output not supported by the provided source material.
Indirect prompt injection
Malicious instructions hidden in content the model reads, such as documents or tool results.

Domain 2: Implement generative AI solutions (18%)

Exam tips

Key terms

Model catalog
The Azure AI Foundry list of models from Microsoft, OpenAI and partners, with model cards and deployment options.
Reasoning model
A model that works through a problem internally before answering, improving accuracy on complex tasks at higher latency and cost.
Embedding model
A model that converts text into a numeric vector representing its meaning, used for search and similarity.
Context window
The maximum number of tokens a model can process across input and output in one request.
System message
The instruction message that sets the model's behavior, rules and format for the whole conversation.
Stateless API
An API that keeps no memory between calls, so the client must resend context each time.
finish_reason
A field explaining why generation stopped: stop, length, content_filter or tool_calls.
Streaming
Returning generated tokens incrementally as they are produced instead of in one final response.
Temperature
A parameter from 0 to 2 that controls randomness; lower values give more deterministic output.
top_p
Nucleus sampling: only tokens within the top cumulative probability p are considered.
Stop sequence
A string that ends generation when the model produces it.
Frequency penalty
A parameter that reduces the likelihood of tokens in proportion to how often they have already appeared.
Few-shot prompting
Including example inputs and desired outputs in the prompt to show the model the task and format.
Zero-shot prompting
Asking the model to perform a task with instructions only and no examples.
Chain-of-thought
Prompting the model to reason step by step before giving its final answer.
Delimiter
Markers such as triple quotes or tags that separate instructions from the content being processed.
RAG
Retrieval augmented generation: retrieving relevant data at query time and adding it to the prompt to ground the answer.
Grounding
Supplying the model with source content it should base its answer on.
Ingestion
The offline process of extracting, chunking, embedding and indexing documents for retrieval.
Citation
A reference in the answer to the source document or chunk that supported it.
Embedding
A numeric vector representation of text in which similar meanings are close together.
Cosine similarity
A measure of how similar two vectors are based on the angle between them.
Chunking
Splitting documents into smaller passages, often with overlap, before embedding and indexing.
HNSW
An approximate nearest neighbor algorithm that finds similar vectors quickly in large indexes.
Flow
A directed graph of tool nodes that defines a generative AI pipeline in prompt flow.
Variant
An alternative version of an LLM node's prompt or parameters, run side by side for comparison.
Connection
A stored, reusable configuration of an external resource's endpoint and credentials used by flow nodes.
Evaluation flow
A flow that scores another flow's outputs against ground truth or quality criteria.
Groundedness
An evaluation metric scoring whether the answer's claims are supported by the provided context.
Relevance
An evaluation metric scoring whether the answer addresses the user's question.
Ground truth
The expected correct answer for a test input, used by similarity and overlap metrics.
Red teaming
Deliberately attacking a system with adversarial inputs to find safety weaknesses before real users do.
Fine-tuning
Further training a base model on your own examples to produce a customized model version.
JSONL
A file format with one JSON object per line, used for fine-tuning training data.
Epoch
One full pass of training over the whole training dataset.
Validation file
Held-out examples used during fine-tuning to measure how well the model generalizes.
Image generation model
A model such as DALL-E 3 that creates new images from a text prompt.
Multimodal model
A model that accepts more than one type of input, such as text and images, in the same request.
Revised prompt
The expanded prompt DALL-E 3 actually used, returned alongside the generated image.
Detail setting
An option that controls how much of an input image the model processes, trading cost for accuracy.
Prompt tokens
Input tokens in a request, including system message, history and retrieved context.
Exponential backoff
Retrying failed requests with increasing wait times to avoid overwhelming a service.
Tracing
Recording each step of a request with inputs, outputs, timing and token use for debugging.
AI gateway
A layer such as Azure API Management in front of model deployments that balances load, enforces limits and logs usage.

Domain 3: Implement an agentic solution (8%)

Exam tips

Key terms

AI agent
Software that uses a model to plan and take actions through tools in a loop to achieve a goal.
Tool
A capability an agent can invoke, such as a function, API, search index or code interpreter.
Instructions
The agent's system-level guidance defining its role, goals, rules and behavior.
Thread
Stored conversation state, including messages and tool results, that an agent works from across turns.
Agent
In Agent Service, a definition combining a model deployment, instructions and tools.
Run
One activation of an agent on a thread that processes messages and tool calls and appends replies.
requires_action
A run status meaning the app must execute requested function calls and submit their outputs.
Run step
An individual action within a run, such as a tool call or message creation, used for inspection and debugging.
File search
A built-in agent tool that indexes uploaded files in a vector store and retrieves passages with citations.
Code interpreter
A built-in agent tool that writes and runs Python in a sandbox for calculations, data processing and charts.
OpenAPI tool
A tool that lets an agent call a REST API described by an OpenAPI specification.
Vector store
The managed store of chunked, embedded files that file search queries.
Function calling
A model feature that returns a structured request to call a described function with arguments, which the app executes.
JSON schema
A standard for describing the structure, types and required fields of JSON data, used for function parameters.
tool_call_id
The identifier linking a tool result message to the model's specific tool call request.
tool_choice
A parameter that lets the model choose tools automatically, forces a specific one or disables them.
Semantic Kernel
Microsoft's open-source SDK for integrating AI models, plugins and agents into C#, Python and Java apps.
Kernel
The Semantic Kernel container holding AI service connections and plugins.
Plugin
A group of functions, described for the model, that Semantic Kernel exposes as callable tools.
AutoGen
A Microsoft Research open-source framework for building conversations between multiple agents.
Orchestration
Coordinating multiple agents so their work combines to complete a task.
Handoff
A pattern where one agent transfers control of a conversation to a more suitable agent.
Connected agents
An Agent Service feature that lets a main agent call specialist agents as tools.
Maker-checker
A pattern where one agent produces work and another reviews it, iterating until accepted.
Least privilege
Granting an agent and its tools only the minimum permissions needed for their task.
Human-in-the-loop
Requiring a person to review or approve an AI action before it takes effect.
Prompt injection
An attack that inserts instructions into model input to override the intended behavior.
OpenTelemetry
An open standard for traces, metrics and logs, used to trace agent runs into Application Insights.
Project endpoint
The Foundry project URL that apps use with the SDK to call Agent Service.
Thread mapping
Storing each user's thread ID so their conversation can continue across requests.
Quality gate
An automated evaluation that must pass before a change is promoted to the next environment.
Standard setup
An Agent Service configuration that uses your own storage, search and networking resources for agent data.

Domain 4: Implement computer vision solutions (13%)

Exam tips

Key terms

Caption
An Image Analysis feature that returns one sentence describing the whole image.
Dense captions
An Image Analysis feature that returns captions with bounding boxes for multiple regions of an image.
Tags
Single-word labels for content in an image, each with a confidence score.
Smart crops
Suggested crop regions that preserve the most important part of an image for given aspect ratios.
OCR
Optical character recognition: extracting machine-readable text from images of text.
Read feature
The Azure AI Vision capability that extracts printed and handwritten text with positions and confidence.
Bounding polygon
The corner points outlining where a line or word appears in the image, including rotated text.
Line and word
The levels at which Read returns text, each with position; words also include confidence scores.
Multiclass classification
A classification type where each image gets exactly one tag.
Multilabel classification
A classification type where each image can have any number of tags.
Object detection
A model type that returns a tag and bounding box for each instance of an object in an image.
Compact domain
A Custom Vision base model that produces smaller, exportable models for offline and edge use.
Iteration
A versioned model produced by one Custom Vision training run.
Precision
The fraction of the model's positive predictions for a tag that were correct.
Recall
The fraction of actual instances of a tag that the model found.
mAP
Mean average precision: the average of per-tag average precision, summarizing detection quality.
Publish name
The name under which an iteration is published to a prediction resource and called by apps.
Prediction resource
The Custom Vision resource that hosts published iterations and serves prediction requests.
Prediction-Key
The HTTP header carrying the prediction resource's key when calling a published model.
Export
Downloading a compact-domain model in a format such as ONNX, TensorFlow, Core ML or a Docker container for offline use.
Face detection
Finding faces in an image and returning their locations, with optional landmarks and attributes.
Face verification
A one-to-one check of whether two faces belong to the same person; a Limited Access feature.
Face identification
A one-to-many search matching a face against enrolled people; a Limited Access feature.
Limited Access
Microsoft's approval process required before using sensitive AI capabilities such as face recognition.
Video Indexer
An Azure AI service that extracts time-coded audio and visual insights from video.
Insights
The JSON output of indexing, such as transcripts, OCR, labels, keywords, faces and scenes, each with timestamps.
Widgets
Embeddable player and insights components for showing indexed video in your own web app.
Keyframe
A representative frame chosen from a shot, used for thumbnails and visual analysis.
Prebuilt model
A model trained by Microsoft that works without your training data.
Custom model
A model trained on your own labeled data to recognize your specific categories.
Multimodal model
A model that accepts images with text and can reason and answer in natural language.
Edge deployment
Running a model on a local device rather than calling a cloud endpoint.

Domain 5: Implement natural language processing solutions (18%)

Exam tips

Key terms

Language detection
Identifying the language of text, returned as a name, ISO code and confidence score.
Named entity recognition
Finding entities in text and classifying them into categories such as Person, Location and Organization.
Entity linking
Disambiguating an entity by linking it to a specific knowledge base entry.
Opinion mining
Aspect-based sentiment that connects sentiment to specific targets mentioned in text.
PII
Personally identifiable information: data that can identify a person, such as names, phone numbers or ID numbers.
Redaction
Masking or replacing sensitive values in text so they are not exposed.
redactedText
The PII detection output field containing the input text with detected entities masked.
PHI
Protected health information, detected by setting the PII domain to phi.
Text translation
Real-time translation of strings through the Translator translate operation.
Document translation
Asynchronous translation of whole files in Blob Storage that preserves formatting.
Custom Translator
A service for training translation models on your own parallel documents, used by category ID.
Transliteration
Converting text from one script to another without translating its meaning.
SpeechConfig
The Speech SDK object holding credentials, region or endpoint, and settings such as recognition language.
recognize_once_async
A Speech SDK method that recognizes a single utterance, ending at a pause.
Continuous recognition
Event-driven recognition of long audio, with interim and final results, until stopped.
Batch transcription
An asynchronous REST API for transcribing large sets of stored audio files.
Neural voice
A prebuilt synthetic voice generated by deep neural networks for natural-sounding speech.
SSML
Speech Synthesis Markup Language, XML that controls voice, pronunciation, pauses, rate and style.
prosody
The SSML element that adjusts speaking rate, pitch and volume.
say-as
The SSML element that controls how values like dates, numbers and letters are read.
SpeechTranslationConfig
The Speech SDK configuration for translating speech, with source language and target languages.
TranslationRecognizer
The Speech SDK recognizer that returns recognized text and its translations.
Intent recognition
Determining the purpose of an utterance and extracting its entities.
Keyword recognition
On-device detection of a wake word using a custom keyword model before full recognition starts.
Intent
The goal or action a user expresses in an utterance, such as BookFlight.
Entity
A piece of information in an utterance the app needs, such as a date or destination.
Utterance
An example phrase a user might say, used to train and test the model.
None intent
The built-in intent that captures utterances outside the app's scope.
Project
A custom question answering knowledge base of question and answer pairs built from sources.
Follow-up prompt
A link from an answer to related pairs, enabling multi-turn conversations.
Synonyms
Project-level word alternatives treated as equivalent when matching questions.
Confidence threshold
The minimum score an answer needs to be returned; below it the default answer is used.
Single-label classification
Custom text classification where each document receives exactly one class.
Multi-label classification
Custom text classification where each document can receive zero or more classes.
Custom NER
A trained model that extracts your own entity types from unstructured documents.
F1 score
The harmonic mean of precision and recall, a single measure of model accuracy.
Custom speech
A speech to text model adapted with your text and audio data for better accuracy in your domain.
Word error rate (WER)
The share of words substituted, deleted or inserted compared with a reference transcript.
Custom neural voice
A Limited Access feature that trains a synthetic voice resembling a specific consenting voice talent.
Phrase list
A runtime list of words that boosts their recognition without training a custom model.

Domain 6: Implement knowledge mining and information extraction solutions (19%)

Exam tips

Key terms

Index
The searchable collection of documents and its schema of fields in Azure AI Search.
Indexer
A crawler that reads a data source, cracks documents, maps fields, runs skillsets and loads the index.
Replica
A copy of the index that serves queries, added for throughput and availability.
Partition
A unit of storage and indexing capacity, added for larger indexes and faster indexing.
Key field
The single Edm.String field that uniquely identifies each document in an index.
Facetable
A field attribute that enables counts of documents per value for faceted navigation.
Analyzer
A component that tokenizes and normalizes text for full-text search, such as a language analyzer.
Suggester
An index definition that enables autocomplete and search-as-you-type suggestions on chosen fields.
Skillset
A collection of AI enrichment skills that an indexer runs during indexing.
Document cracking
The indexer step that opens files and extracts text, metadata and images.
Enrichment tree
The in-memory structure holding a document's original and enriched data as skills run.
Output field mapping
A mapping that copies enriched values from the enrichment tree into index fields.
Custom skill
A skill that calls your own code, such as an Azure Function, during AI enrichment.
WebApiSkill
The skill type that calls an external HTTPS endpoint with a defined JSON interface.
recordId
The identifier that links each input record to its output in the custom skill interface.
values array
The top-level array of records sent to and returned from a custom skill.
queryType
The query parser: simple (default), full (Lucene) or semantic.
$filter
An OData expression that restricts results by filterable field values without affecting scoring.
Fuzzy search
A full Lucene query that matches terms with small spelling differences, written with ~.
Facets
Counts of matching documents per field value or range, used for faceted navigation.
Vector search
Finding documents whose embedding vectors are closest to the query's vector.
Hybrid search
Combining keyword and vector queries in one request, merged with Reciprocal Rank Fusion.
Reciprocal Rank Fusion
A method that merges ranked lists by scoring documents on their rank positions in each list.
Semantic ranker
A second-stage model that re-ranks top results by meaning and can return captions and answers.
Knowledge store
Azure Storage output of a skillset's enriched data, saved for uses beyond search.
Table projection
A projection that writes enriched data as rows in Azure Table Storage.
Object projection
A projection that writes enriched data as JSON documents in Blob Storage.
Shaper skill
A utility skill that builds a custom data shape from the enrichment tree for projection.
Read model
The Document Intelligence model that extracts text lines and words from documents.
Layout model
The model that extracts text plus structure: paragraphs, tables, selection marks and figures.
Prebuilt model
A ready-to-use model that extracts named fields from a common document type such as invoices or receipts.
Confidence score
A value from 0 to 1 indicating how certain the model is about an extracted field.
Custom template model
A custom extraction model for documents with a consistent, fixed layout.
Custom neural model
A deep learning custom extraction model that handles documents with varying layouts.
Custom classifier
A model that identifies document types, and splits combined files, before extraction.
Composed model
Several custom models grouped under one model ID that routes each document to the best match.
Content Understanding
An Azure AI service that extracts structured, schema-defined output from documents, images, audio and video.
Field schema
The list of named, typed and described fields an analyzer should return.
Generate field
A schema field whose value the model produces from the content, such as a summary.
Study Azure AI Engineer for free
Lessons, quizzes, exam simulations and hands-on labs.
Open the Azure AI Engineer study plan