All certifications / Azure AI Engineer / Cheat sheet
Azure AI Engineer AI-102 cheat sheet
Domain 1: Plan and manage an Azure AI solution (24%)
Exam tips
- Match the output the scenario asks for: named fields from forms means Document Intelligence, plain text lines means Vision Read, harmful-content scores means Content Safety, and generated prose means Azure OpenAI.
- If the question stresses one endpoint and one bill for several prebuilt services, pick the multi-service resource. If it stresses a Free tier or separate keys per service, pick single-service resources. Shared governance with separate workspaces means a hub with projects.
- Global means any region, Data Zone means within the EU or US zone, Standard means the resource's region, and Provisioned means reserved capacity with predictable latency. A 429 error points to quota or rate limits.
- In a template question, look at kind to identify the service and sku.name for the tier. customSubDomainName is the property tied to Entra ID authentication.
- Know the status codes: 401 credential problem, 403 permission or network block, 404 wrong path or deployment name, 429 rate limit, 202 asynchronous operation started.
- Reader is a management-plane role and cannot call AI APIs. For keyless calls choose a managed identity plus Cognitive Services User (or OpenAI User for Azure OpenAI). Token auth needs a custom subdomain.
- Rotation without downtime is always switch to the other key, then regenerate. Keeping a resource off the internet is a private endpoint plus disabling public network access.
- Know the three required container parameters: ApiKey, Billing and Eula. No internet at all means disconnected containers, which need Microsoft approval and a commitment plan.
- Metrics are automatic and good for alerts; logs need a diagnostic setting and go to Log Analytics for KQL. Every alert notifies through an action group.
- Map the scenario to the principle: explaining AI use and limits is transparency, equal treatment of groups is fairness, accessibility is inclusiveness, human ownership of outcomes is accountability. Specific forbidden words mean a blocklist.
- 400 with content_filter means the prompt was blocked; finish_reason content_filter means the output was blocked. Jailbreak in the user message is a user prompt attack; instructions hidden in data are a document attack.
Key terms
- Azure AI services
- Microsoft's family of prebuilt AI APIs for vision, language, speech, translation, documents and content safety, called over REST or SDKs.
- Azure OpenAI
- Azure-hosted OpenAI models for chat, reasoning, embeddings and image generation, deployed and managed through Azure AI Foundry.
- Document Intelligence
- A service that extracts text, structure and named fields from documents using prebuilt or custom models.
- Content Safety
- A service that detects harmful text and images across hate, sexual, violence and self-harm categories.
- Single-service resource
- An Azure resource for one AI service, with its own endpoint, keys, pricing tier and bill.
- Multi-service resource
- An Azure AI services resource that exposes many AI services behind one endpoint and key pair with combined billing.
- Hub
- A shared Azure AI Foundry container, built on Azure Machine Learning, that holds connections, security and compute for several projects.
- Project
- A Foundry workspace where you deploy models, build agents, evaluate and manage the assets for one app or team.
- Deployment
- A named instance of a model version inside an Azure OpenAI or Foundry resource, called by name from code.
- Tokens per minute (TPM)
- The rate limit and quota unit for pay-as-you-go deployments.
- Provisioned throughput unit (PTU)
- A unit of reserved model processing capacity for provisioned deployments.
- Data zone
- A group of regions, such as the EU or US, within which Data Zone deployments process requests.
- kind
- The template property that sets which AI service a Microsoft.CognitiveServices/accounts resource provides.
- Bicep
- A domain-specific language for Azure infrastructure as code that compiles to ARM JSON templates.
- customSubDomainName
- The property that gives a resource a unique endpoint name, required for Entra ID authentication and private endpoints.
- disableLocalAuth
- A property that turns off key-based authentication so only Microsoft Entra ID tokens work.
- Endpoint
- The base URL of an Azure AI resource that all API calls are sent to.
- Ocp-Apim-Subscription-Key
- The HTTP header that carries a resource key for most Azure AI services.
- DefaultAzureCredential
- An Azure Identity class that finds a Microsoft Entra ID credential automatically from the environment, managed identity or developer sign-in.
- Operation-Location
- The header returned by asynchronous operations that points to the URL you poll for the result.
- Managed identity
- An Entra ID identity for an Azure resource whose credentials Azure creates and rotates automatically.
- Cognitive Services User
- An RBAC role that lets an identity call Azure AI services data-plane APIs.
- Cognitive Services OpenAI User
- An RBAC role that allows inference calls to Azure OpenAI deployments without management rights.
- Custom subdomain
- A unique resource-specific endpoint name required for Entra ID token authentication.
- Azure Key Vault
- A managed service for storing and auditing secrets, keys and certificates.
- Key rotation
- Replacing keys regularly, using the two-key pattern so clients never lose access.
- Private endpoint
- A network interface with a private IP in your virtual network that connects privately to an Azure resource.
- Customer-managed key
- An encryption key you own in Key Vault used to encrypt a service's data at rest instead of a Microsoft-managed key.
- Connected container
- A container that processes data locally but reports usage to Azure for billing.
- Disconnected container
- A container approved for fully offline use under a commitment plan, with no online usage reporting.
- Billing setting
- The container startup value that holds the endpoint of the Azure resource charged for usage.
- Eula=accept
- The required startup argument confirming acceptance of the container's license terms.
- Diagnostic setting
- A resource configuration that routes logs and metrics to Log Analytics, storage or an event hub.
- Action group
- A reusable set of notifications and automated actions triggered by alerts.
- Commitment tier
- A pricing plan with a fixed monthly fee for a set volume of usage at a discounted rate.
- Budget
- A Cost Management threshold on actual or forecasted spend that sends alerts.
- Harm categories
- The four Content Safety classifications: hate, sexual, violence and self-harm.
- Severity level
- A score showing how harmful content is within a category, which your app compares to a threshold.
- Blocklist
- A custom list of terms that Content Safety flags in addition to its harm classifiers.
- Transparency
- The responsible AI principle that people should know they are using AI and understand its capabilities and limits.
- Content filter
- The configurable input and output moderation applied to every Azure OpenAI deployment.
- Prompt shields
- A Content Safety feature that detects user prompt attacks (jailbreaks) and document attacks (indirect prompt injection).
- Groundedness detection
- A check that flags model output not supported by the provided source material.
- Indirect prompt injection
- Malicious instructions hidden in content the model reads, such as documents or tool results.
Domain 2: Implement generative AI solutions (18%)
Exam tips
- Vectors for search means an embedding model; new pictures means an image model; hard multistep logic suggests a reasoning model; everything conversational is a chat model. Code calls the deployment name, not the model name.
- Rules and persona go in the system message. Memory in a chat app comes from resending prior messages, not from the deployment. finish_reason length means raise max tokens.
- Consistency means lower temperature. Cut-off answers with finish_reason length mean raise max tokens. Repetition means frequency or presence penalty. Change temperature or top_p, not both.
- When the scenario is about getting an exact format or style, the answer is usually few-shot examples or structured output. Instructions and rules belong in the system message; content should be clearly delimited.
- Up-to-date facts, private documents and citations point to RAG, not fine-tuning. If answers are wrong because the right passage is missing, fix retrieval (hybrid search, semantic ranking, better chunking) before changing the model.
- The vector field's dimensions must match the embedding model, and queries must be embedded with the same model as the documents. Chunk overlap prevents losing context at boundaries.
- Comparing prompt versions in one node means variants; storing endpoints and keys means connections; a conversational flow with history is a chat flow; publishing it as an API means a managed online endpoint.
- Answers not supported by retrieved context point to groundedness; off-topic answers point to relevance; poor readability points to fluency or coherence. Overlap metrics such as F1 and BLEU need ground truth.
- Facts that change and need citations: RAG. Consistent style or format, or shorter prompts on a narrow task: fine-tuning. Try prompt engineering first in every case.
- Creating pictures means an image generation model; asking questions about a picture means sending it as an image part to a vision-enabled chat model. For high-volume structured extraction, prefer Document Intelligence or Vision.
- 429 means rate limits: back off and retry, add quota or deployments, or use provisioned throughput. Finding which step caused a slow or bad answer means tracing with Application Insights.
Key terms
- Model catalog
- The Azure AI Foundry list of models from Microsoft, OpenAI and partners, with model cards and deployment options.
- Reasoning model
- A model that works through a problem internally before answering, improving accuracy on complex tasks at higher latency and cost.
- Embedding model
- A model that converts text into a numeric vector representing its meaning, used for search and similarity.
- Context window
- The maximum number of tokens a model can process across input and output in one request.
- System message
- The instruction message that sets the model's behavior, rules and format for the whole conversation.
- Stateless API
- An API that keeps no memory between calls, so the client must resend context each time.
- finish_reason
- A field explaining why generation stopped: stop, length, content_filter or tool_calls.
- Streaming
- Returning generated tokens incrementally as they are produced instead of in one final response.
- Temperature
- A parameter from 0 to 2 that controls randomness; lower values give more deterministic output.
- top_p
- Nucleus sampling: only tokens within the top cumulative probability p are considered.
- Stop sequence
- A string that ends generation when the model produces it.
- Frequency penalty
- A parameter that reduces the likelihood of tokens in proportion to how often they have already appeared.
- Few-shot prompting
- Including example inputs and desired outputs in the prompt to show the model the task and format.
- Zero-shot prompting
- Asking the model to perform a task with instructions only and no examples.
- Chain-of-thought
- Prompting the model to reason step by step before giving its final answer.
- Delimiter
- Markers such as triple quotes or tags that separate instructions from the content being processed.
- RAG
- Retrieval augmented generation: retrieving relevant data at query time and adding it to the prompt to ground the answer.
- Grounding
- Supplying the model with source content it should base its answer on.
- Ingestion
- The offline process of extracting, chunking, embedding and indexing documents for retrieval.
- Citation
- A reference in the answer to the source document or chunk that supported it.
- Embedding
- A numeric vector representation of text in which similar meanings are close together.
- Cosine similarity
- A measure of how similar two vectors are based on the angle between them.
- Chunking
- Splitting documents into smaller passages, often with overlap, before embedding and indexing.
- HNSW
- An approximate nearest neighbor algorithm that finds similar vectors quickly in large indexes.
- Flow
- A directed graph of tool nodes that defines a generative AI pipeline in prompt flow.
- Variant
- An alternative version of an LLM node's prompt or parameters, run side by side for comparison.
- Connection
- A stored, reusable configuration of an external resource's endpoint and credentials used by flow nodes.
- Evaluation flow
- A flow that scores another flow's outputs against ground truth or quality criteria.
- Groundedness
- An evaluation metric scoring whether the answer's claims are supported by the provided context.
- Relevance
- An evaluation metric scoring whether the answer addresses the user's question.
- Ground truth
- The expected correct answer for a test input, used by similarity and overlap metrics.
- Red teaming
- Deliberately attacking a system with adversarial inputs to find safety weaknesses before real users do.
- Fine-tuning
- Further training a base model on your own examples to produce a customized model version.
- JSONL
- A file format with one JSON object per line, used for fine-tuning training data.
- Epoch
- One full pass of training over the whole training dataset.
- Validation file
- Held-out examples used during fine-tuning to measure how well the model generalizes.
- Image generation model
- A model such as DALL-E 3 that creates new images from a text prompt.
- Multimodal model
- A model that accepts more than one type of input, such as text and images, in the same request.
- Revised prompt
- The expanded prompt DALL-E 3 actually used, returned alongside the generated image.
- Detail setting
- An option that controls how much of an input image the model processes, trading cost for accuracy.
- Prompt tokens
- Input tokens in a request, including system message, history and retrieved context.
- Exponential backoff
- Retrying failed requests with increasing wait times to avoid overwhelming a service.
- Tracing
- Recording each step of a request with inputs, outputs, timing and token use for debugging.
- AI gateway
- A layer such as Azure API Management in front of model deployments that balances load, enforces limits and logs usage.
Domain 3: Implement an agentic solution (8%)
Exam tips
- Multi-step tasks that call APIs or take actions suggest an agent. Single-shot text transformations (rewrite, translate, summarize one document) suggest a plain chat completion.
- Conversation history lives in the thread, behavior in the agent, and processing in the run. requires_action means your code must run the function and submit tool outputs.
- Calculations or charts from uploaded data: code interpreter. Answers from uploaded documents: file search. Existing enterprise index: Azure AI Search tool. Existing REST API with a spec: OpenAPI. Your own code in your app: function calling.
- The model only proposes calls; your code executes them and must send results back with the matching tool call ID. Descriptions drive which function the model chooses.
- Exposing your own methods to the model in Semantic Kernel means a plugin with kernel functions. AutoGen is associated with multi-agent conversations. Agent Service is the managed, hosted option.
- A main agent delegating to specialists maps to connected agents or handoff. Agents reviewing each other's work is group chat or maker-checker. Fixed step order is sequential orchestration.
- Risky or irreversible actions need human approval enforced in code. Limit damage with least-privilege tool permissions. Finding out which tool call went wrong needs tracing.
- Apps call hosted agents through the project endpoint with Entra ID, ideally a managed identity. Framework-based agents are deployed like any app (Container Apps, App Service, Functions).
Key terms
- AI agent
- Software that uses a model to plan and take actions through tools in a loop to achieve a goal.
- Tool
- A capability an agent can invoke, such as a function, API, search index or code interpreter.
- Instructions
- The agent's system-level guidance defining its role, goals, rules and behavior.
- Thread
- Stored conversation state, including messages and tool results, that an agent works from across turns.
- Agent
- In Agent Service, a definition combining a model deployment, instructions and tools.
- Run
- One activation of an agent on a thread that processes messages and tool calls and appends replies.
- requires_action
- A run status meaning the app must execute requested function calls and submit their outputs.
- Run step
- An individual action within a run, such as a tool call or message creation, used for inspection and debugging.
- File search
- A built-in agent tool that indexes uploaded files in a vector store and retrieves passages with citations.
- Code interpreter
- A built-in agent tool that writes and runs Python in a sandbox for calculations, data processing and charts.
- OpenAPI tool
- A tool that lets an agent call a REST API described by an OpenAPI specification.
- Vector store
- The managed store of chunked, embedded files that file search queries.
- Function calling
- A model feature that returns a structured request to call a described function with arguments, which the app executes.
- JSON schema
- A standard for describing the structure, types and required fields of JSON data, used for function parameters.
- tool_call_id
- The identifier linking a tool result message to the model's specific tool call request.
- tool_choice
- A parameter that lets the model choose tools automatically, forces a specific one or disables them.
- Semantic Kernel
- Microsoft's open-source SDK for integrating AI models, plugins and agents into C#, Python and Java apps.
- Kernel
- The Semantic Kernel container holding AI service connections and plugins.
- Plugin
- A group of functions, described for the model, that Semantic Kernel exposes as callable tools.
- AutoGen
- A Microsoft Research open-source framework for building conversations between multiple agents.
- Orchestration
- Coordinating multiple agents so their work combines to complete a task.
- Handoff
- A pattern where one agent transfers control of a conversation to a more suitable agent.
- Connected agents
- An Agent Service feature that lets a main agent call specialist agents as tools.
- Maker-checker
- A pattern where one agent produces work and another reviews it, iterating until accepted.
- Least privilege
- Granting an agent and its tools only the minimum permissions needed for their task.
- Human-in-the-loop
- Requiring a person to review or approve an AI action before it takes effect.
- Prompt injection
- An attack that inserts instructions into model input to override the intended behavior.
- OpenTelemetry
- An open standard for traces, metrics and logs, used to trace agent runs into Application Insights.
- Project endpoint
- The Foundry project URL that apps use with the SDK to call Agent Service.
- Thread mapping
- Storing each user's thread ID so their conversation can continue across requests.
- Quality gate
- An automated evaluation that must pass before a change is promoted to the next environment.
- Standard setup
- An Agent Service configuration that uses your own storage, search and networking resources for agent data.
Domain 4: Implement computer vision solutions (13%)
Exam tips
- One sentence for the whole image is caption; sentences with boxes for regions is dense captions; single words are tags; where things are is objects; detecting people is not identifying them.
- Text in photos and scenes points to Vision Read; long documents, tables and form fields point to Document Intelligence. Read returns lines and words with polygons and word confidence.
- Exactly one label per image: multiclass. Several labels per image: multilabel. Location or count: object detection. Offline or edge: compact domain.
- Precision is about how many predictions were right; recall is about how many real cases were found. Raising the threshold trades recall for precision. mAP is the headline metric for object detection.
- Retrained but no change in the app: the new iteration was not published under the name the app calls. Offline use: export, which requires a compact domain.
- Detection and quality attributes are available; identification and verification need Limited Access approval; emotion, age and gender inference were retired.
- Searchable transcripts, on-screen text, topics and scenes from recorded videos point to Video Indexer. Recognizing specific faces in video still requires Limited Access approval.
- General concepts, no training: Image Analysis. Your own categories at scale or offline: Custom Vision. Open-ended questions or reasoning about an image: a multimodal model. Forms: Document Intelligence.
Key terms
- Caption
- An Image Analysis feature that returns one sentence describing the whole image.
- Dense captions
- An Image Analysis feature that returns captions with bounding boxes for multiple regions of an image.
- Tags
- Single-word labels for content in an image, each with a confidence score.
- Smart crops
- Suggested crop regions that preserve the most important part of an image for given aspect ratios.
- OCR
- Optical character recognition: extracting machine-readable text from images of text.
- Read feature
- The Azure AI Vision capability that extracts printed and handwritten text with positions and confidence.
- Bounding polygon
- The corner points outlining where a line or word appears in the image, including rotated text.
- Line and word
- The levels at which Read returns text, each with position; words also include confidence scores.
- Multiclass classification
- A classification type where each image gets exactly one tag.
- Multilabel classification
- A classification type where each image can have any number of tags.
- Object detection
- A model type that returns a tag and bounding box for each instance of an object in an image.
- Compact domain
- A Custom Vision base model that produces smaller, exportable models for offline and edge use.
- Iteration
- A versioned model produced by one Custom Vision training run.
- Precision
- The fraction of the model's positive predictions for a tag that were correct.
- Recall
- The fraction of actual instances of a tag that the model found.
- mAP
- Mean average precision: the average of per-tag average precision, summarizing detection quality.
- Publish name
- The name under which an iteration is published to a prediction resource and called by apps.
- Prediction resource
- The Custom Vision resource that hosts published iterations and serves prediction requests.
- Prediction-Key
- The HTTP header carrying the prediction resource's key when calling a published model.
- Export
- Downloading a compact-domain model in a format such as ONNX, TensorFlow, Core ML or a Docker container for offline use.
- Face detection
- Finding faces in an image and returning their locations, with optional landmarks and attributes.
- Face verification
- A one-to-one check of whether two faces belong to the same person; a Limited Access feature.
- Face identification
- A one-to-many search matching a face against enrolled people; a Limited Access feature.
- Limited Access
- Microsoft's approval process required before using sensitive AI capabilities such as face recognition.
- Video Indexer
- An Azure AI service that extracts time-coded audio and visual insights from video.
- Insights
- The JSON output of indexing, such as transcripts, OCR, labels, keywords, faces and scenes, each with timestamps.
- Widgets
- Embeddable player and insights components for showing indexed video in your own web app.
- Keyframe
- A representative frame chosen from a shot, used for thumbnails and visual analysis.
- Prebuilt model
- A model trained by Microsoft that works without your training data.
- Custom model
- A model trained on your own labeled data to recognize your specific categories.
- Multimodal model
- A model that accepts images with text and can reason and answer in natural language.
- Edge deployment
- Running a model on a local device rather than calling a cloud endpoint.
Domain 5: Implement natural language processing solutions (18%)
Exam tips
- Overall positive or negative is sentiment; which aspect is liked or disliked is opinion mining. Categorizing entities is NER; resolving which real thing an entity refers to is entity linking.
- Mask personal data in text: PII detection with redactedText. Restrict which types are masked with the categories filter. Healthcare data: domain phi.
- Several target languages in one call: multiple to parameters. 401 with a regional key: add the Ocp-Apim-Subscription-Region header. Whole files with formatting: document translation. Domain terminology: Custom Translator with a category ID.
- One short command: recognize once. Long live audio: continuous recognition with events. Thousands of stored files: batch transcription. A Canceled result usually means credential, region or connection problems.
- Speed and pitch: prosody. Pauses: break. Dates, numbers and spelling: say-as. Exact pronunciation: phoneme or a lexicon. Emotion or style: mstts:express-as.
- Live speech into another language: SpeechTranslationConfig with add_target_language and a TranslationRecognizer. A wake word: keyword recognition with a custom keyword model. Understanding commands: CLU intents and entities.
- Out-of-scope input landing in a real intent: add examples to None. A fixed set of values with synonyms: list entity. Patterns like codes: regex entity. Values that vary in context: learned entity.
- Clarifying follow-ups: multi-turn prompts. Equivalent words across the whole project: synonyms. Wrong answers to unrelated questions: raise the confidence threshold and set a default answer. It needs an Azure AI Search resource.
- One category per document: single-label. Any number: multi-label. Your own entity types from long documents: custom NER. Training data lives in a Blob Storage container connected to the Language resource.
- Vocabulary problems: plain text data or a phrase list. Noise or accent problems: audio plus human-labeled transcripts. Custom neural voice needs Limited Access approval and the talent's recorded consent.
Key terms
- Language detection
- Identifying the language of text, returned as a name, ISO code and confidence score.
- Named entity recognition
- Finding entities in text and classifying them into categories such as Person, Location and Organization.
- Entity linking
- Disambiguating an entity by linking it to a specific knowledge base entry.
- Opinion mining
- Aspect-based sentiment that connects sentiment to specific targets mentioned in text.
- PII
- Personally identifiable information: data that can identify a person, such as names, phone numbers or ID numbers.
- Redaction
- Masking or replacing sensitive values in text so they are not exposed.
- redactedText
- The PII detection output field containing the input text with detected entities masked.
- PHI
- Protected health information, detected by setting the PII domain to phi.
- Text translation
- Real-time translation of strings through the Translator translate operation.
- Document translation
- Asynchronous translation of whole files in Blob Storage that preserves formatting.
- Custom Translator
- A service for training translation models on your own parallel documents, used by category ID.
- Transliteration
- Converting text from one script to another without translating its meaning.
- SpeechConfig
- The Speech SDK object holding credentials, region or endpoint, and settings such as recognition language.
- recognize_once_async
- A Speech SDK method that recognizes a single utterance, ending at a pause.
- Continuous recognition
- Event-driven recognition of long audio, with interim and final results, until stopped.
- Batch transcription
- An asynchronous REST API for transcribing large sets of stored audio files.
- Neural voice
- A prebuilt synthetic voice generated by deep neural networks for natural-sounding speech.
- SSML
- Speech Synthesis Markup Language, XML that controls voice, pronunciation, pauses, rate and style.
- prosody
- The SSML element that adjusts speaking rate, pitch and volume.
- say-as
- The SSML element that controls how values like dates, numbers and letters are read.
- SpeechTranslationConfig
- The Speech SDK configuration for translating speech, with source language and target languages.
- TranslationRecognizer
- The Speech SDK recognizer that returns recognized text and its translations.
- Intent recognition
- Determining the purpose of an utterance and extracting its entities.
- Keyword recognition
- On-device detection of a wake word using a custom keyword model before full recognition starts.
- Intent
- The goal or action a user expresses in an utterance, such as BookFlight.
- Entity
- A piece of information in an utterance the app needs, such as a date or destination.
- Utterance
- An example phrase a user might say, used to train and test the model.
- None intent
- The built-in intent that captures utterances outside the app's scope.
- Project
- A custom question answering knowledge base of question and answer pairs built from sources.
- Follow-up prompt
- A link from an answer to related pairs, enabling multi-turn conversations.
- Synonyms
- Project-level word alternatives treated as equivalent when matching questions.
- Confidence threshold
- The minimum score an answer needs to be returned; below it the default answer is used.
- Single-label classification
- Custom text classification where each document receives exactly one class.
- Multi-label classification
- Custom text classification where each document can receive zero or more classes.
- Custom NER
- A trained model that extracts your own entity types from unstructured documents.
- F1 score
- The harmonic mean of precision and recall, a single measure of model accuracy.
- Custom speech
- A speech to text model adapted with your text and audio data for better accuracy in your domain.
- Word error rate (WER)
- The share of words substituted, deleted or inserted compared with a reference transcript.
- Custom neural voice
- A Limited Access feature that trains a synthetic voice resembling a specific consenting voice talent.
- Phrase list
- A runtime list of words that boosts their recognition without training a custom model.
Domain 6: Implement knowledge mining and information extraction solutions (19%)
Exam tips
- Slow queries or availability: add replicas. Bigger index or slow indexing: add partitions. Pull from Azure data on a schedule: indexer. Plan the tier up front: downgrades are not possible and in-place upgrades are limited to eligible services.
- Map the query feature to the attribute: $filter needs filterable, $orderby needs sortable, facets need facetable, full-text needs searchable, returned in results needs retrievable. Autocomplete needs a suggester.
- Scanned text needs imageAction generateNormalizedImages plus the OCR skill. Enriching more than a handful of documents needs an attached Azure AI services resource. Enriched values reach the index through output field mappings.
- Your own model or lookup during indexing means a custom Web API skill. The endpoint must accept and return a values array with matching recordIds and a data object.
- Fuzzy, proximity, boosting, regex and fielded queries need queryType full. Filters are OData with word operators (eq, ge, and), not symbols like == or &&.
- Exact codes plus meaning: hybrid search. Better ordering of the top results and extracted answers: semantic ranking with a semantic configuration and queryType semantic. Queries must use the same embedding model as the documents.
- Analytics in Power BI: table projections. Hierarchical JSON for other systems: object projections. Extracted images: file projections. Related tables that must join: same projection group.
- Text only: read. Tables, checkboxes and structure without named fields: layout. Named fields from common documents: the matching prebuilt model. 202 plus Operation-Location means poll for the result.
- Fixed layout: template. Varying layouts from many sources: neural. Mixed document types in one stream: custom classifier. One model ID for several custom models: composed model.
- One service with a field schema across documents, images, audio and video points to Content Understanding. The analyzer holds the schema; extract, generate and classify are the field methods.
Key terms
- Index
- The searchable collection of documents and its schema of fields in Azure AI Search.
- Indexer
- A crawler that reads a data source, cracks documents, maps fields, runs skillsets and loads the index.
- Replica
- A copy of the index that serves queries, added for throughput and availability.
- Partition
- A unit of storage and indexing capacity, added for larger indexes and faster indexing.
- Key field
- The single Edm.String field that uniquely identifies each document in an index.
- Facetable
- A field attribute that enables counts of documents per value for faceted navigation.
- Analyzer
- A component that tokenizes and normalizes text for full-text search, such as a language analyzer.
- Suggester
- An index definition that enables autocomplete and search-as-you-type suggestions on chosen fields.
- Skillset
- A collection of AI enrichment skills that an indexer runs during indexing.
- Document cracking
- The indexer step that opens files and extracts text, metadata and images.
- Enrichment tree
- The in-memory structure holding a document's original and enriched data as skills run.
- Output field mapping
- A mapping that copies enriched values from the enrichment tree into index fields.
- Custom skill
- A skill that calls your own code, such as an Azure Function, during AI enrichment.
- WebApiSkill
- The skill type that calls an external HTTPS endpoint with a defined JSON interface.
- recordId
- The identifier that links each input record to its output in the custom skill interface.
- values array
- The top-level array of records sent to and returned from a custom skill.
- queryType
- The query parser: simple (default), full (Lucene) or semantic.
- $filter
- An OData expression that restricts results by filterable field values without affecting scoring.
- Fuzzy search
- A full Lucene query that matches terms with small spelling differences, written with ~.
- Facets
- Counts of matching documents per field value or range, used for faceted navigation.
- Vector search
- Finding documents whose embedding vectors are closest to the query's vector.
- Hybrid search
- Combining keyword and vector queries in one request, merged with Reciprocal Rank Fusion.
- Reciprocal Rank Fusion
- A method that merges ranked lists by scoring documents on their rank positions in each list.
- Semantic ranker
- A second-stage model that re-ranks top results by meaning and can return captions and answers.
- Knowledge store
- Azure Storage output of a skillset's enriched data, saved for uses beyond search.
- Table projection
- A projection that writes enriched data as rows in Azure Table Storage.
- Object projection
- A projection that writes enriched data as JSON documents in Blob Storage.
- Shaper skill
- A utility skill that builds a custom data shape from the enrichment tree for projection.
- Read model
- The Document Intelligence model that extracts text lines and words from documents.
- Layout model
- The model that extracts text plus structure: paragraphs, tables, selection marks and figures.
- Prebuilt model
- A ready-to-use model that extracts named fields from a common document type such as invoices or receipts.
- Confidence score
- A value from 0 to 1 indicating how certain the model is about an extracted field.
- Custom template model
- A custom extraction model for documents with a consistent, fixed layout.
- Custom neural model
- A deep learning custom extraction model that handles documents with varying layouts.
- Custom classifier
- A model that identifies document types, and splits combined files, before extraction.
- Composed model
- Several custom models grouped under one model ID that routes each document to the best match.
- Content Understanding
- An Azure AI service that extracts structured, schema-defined output from documents, images, audio and video.
- Field schema
- The list of named, typed and described fields an analyzer should return.
- Generate field
- A schema field whose value the model produces from the content, such as a summary.
Study Azure AI Engineer for free
Lessons, quizzes, exam simulations and hands-on labs.
Open the Azure AI Engineer study planLessons, quizzes, exam simulations and hands-on labs.