Free live cohort on Google Meet — register your interest →
AWS Certified AI Practitioner · Full practice exam

Full Practice Exam

65 questions across all 5 domains, sampled at the official exam weights. Every answer is graded and explained.

Domain 1
13 questions Fundamentals of AI and ML
Domain 2
16 questions Fundamentals of Generative AI
Domain 3
18 questions Applications of Foundation Models
Domain 4
9 questions Guidelines for Responsible AI
Domain 5
9 questions Security, Compliance, and Governance for AI Solutions

Chapter 18 — Full Practice Exam and Exam-Day Strategy

Certification Blueprint

Field Coverage
Exam AWS Certified AI Practitioner (AIF-C01), exam guide v1.1 (2026-04-30)
Primary domain All five, sampled at 20 / 24 / 28 / 14 / 14
Task statements Every task statement, reached through the chapters that teach it
Exam facts 65 questions presented · 50 scored · 15 unscored trial · scaled 100-1000 · pass at 700 · compensatory
Question types Multiple choice · multiple response · ordering · matching
Exam-guide status Verified live 2026-08-07. See course-plan.md § Certification Provenance

What This Chapter Covers

Everything and nothing. There is no new subject matter here — every question on the practice exam is answerable from Chapters 01-17. What is new is the form: domains mixed, unlabelled, in sequence, with no chapter heading to tell you which mental model to reach for.

That difference matters more than it sounds. Studying a chapter primes you: you know the answer is about governance because the chapter is about governance. The exam removes that scaffolding entirely, and a candidate who has only ever been tested with the scaffolding in place has not yet been tested.

This chapter therefore teaches three things the earlier chapters could not:

  1. Retrieval under mixing — recognising which domain a question belongs to before answering it.
  2. Triage under time pressure — deciding fast which questions to fight and which to bank.
  3. Reading your own score — converting a raw mark into a specific list of chapters.

Core Mental Model

The 65 presented questions, of which 50 are scored and 15 are unscored trial items, dividing by published domain weight into 13 questions for Domain 1 fundamentals of AI and ML sourced from chapters 01 to 03, 16 for Domain 2 fundamentals of generative AI from chapters 04 to 07, 18 for Domain 3 applications of foundation models from chapters 08 to 12, 9 for Domain 4 responsible AI from chapters 13 to 15, and 9 for Domain 5 security and governance from chapters 16 to 17, with all five domains converging on a single scaled score of 100 to 1000 passing at 700 under a compensatory model that sets no per-domain pass bar

What to remember from this diagram: the arrows converge. Five domains, five different bodies of knowledge, and exactly one number at the end. There is no per-domain pass bar, so the exam does not care where your marks come from. That single fact drives every strategic decision in this chapter.

Two consequences follow immediately, and both are counter-intuitive:

  • A weak domain is survivable. Domains 4 and 5 carry 14% each. Answering every one of those 18 questions wrong and everything else right still passes comfortably. Panic about a weak domain is usually more expensive than the weak domain.
  • Domain 3 is worth more than Domains 4 and 5 combined. 28% against 28% — the same. But Domain 3 is one body of knowledge and the other two are two. Marks per unit of revision are highest in Domain 3, and it is where a limited revision budget should go first.

Detailed Concept Explanations

The 15 unscored questions, and why they change nothing

The exam presents 65 questions and scores 50. The other 15 are trial items being calibrated for future papers, and they are not identifiable — no marker, no section, no pattern.

The temptation is to treat this as a licence to relax: "one in four doesn't count." Resist it. You cannot tell which, so the only rational strategy is to treat all 65 as scored. What the fact genuinely buys you is emotional: when you meet a question that seems impossibly obscure, or oddly-worded, or about something you are certain was never in the guide, there is a real chance it is a trial item. Do not let it shake you. Answer it, flag it, move on. Candidates lose marks on the questions after a demoralising one far more often than on the demoralising one itself.

Compensatory scoring, stated precisely

Your result is one scaled score from 100 to 1000, with 700 to pass. It is not a percentage, and the scaling is not published, so do not try to convert your raw practice mark into a scaled score — the arithmetic does not exist publicly and anyone who offers you the conversion is guessing.

Compensatory means the domains pool. There is no rule of the form "you must reach X% in Domain 4." This is the single most actionable fact about the exam's structure, and it produces the triage rule below.

The four question types

Type What it asks The trap
Multiple choice One correct response, three distractors The longest option looks most complete
Multiple response Two or more correct from five or more Partial credit does not exist — half right scores zero
Ordering Place steps into the correct sequence Anchor an end of the sequence first, not the middle
Matching Pair items across two lists Do the pairs you are certain of first; the rest constrain each other

Multiple response deserves specific respect. Selecting one of two correct answers scores exactly the same as selecting two wrong ones: nothing. When a question says select two, verify you have selected two before moving on — an unselected second answer is the most common unforced error on this exam, and it is entirely mechanical.

Reading a question in one pass

The strongest single habit is to identify the constraint before looking at the options. Almost every scenario question contains one clause that decides the answer — "with minimal infrastructure to manage", "for every day of the preceding quarter", "which identity", "before a model is built". Find it, then read the options against it.

Read the options after the constraint, not before. Reading options first is how a plausible, well-written distractor gets to establish itself as the answer before you have worked out what the question needs.

Decision Rules and Exam Signals

Signal in the stem What it selects Chapter
"Guaranteed identical output" / "auditable to a fixed rule" Reject AI entirely 02
"Without labels" / "groups we have not defined" Unsupervised, clustering 01, 02
"Minimal infrastructure to manage" The managed service, not the capable one 07
"For every day of" / "state across a period" AWS Config 17
"Which identity" / "who changed" AWS CloudTrail 17
"Before a model is built" SageMaker Clarify, not Model Monitor 14
"Must cite the source passage" RAG, not fine-tuning 09, 11
"Consistent tone or format" Fine-tuning, not RAG 11
"Documents change several times a day" RAG, not fine-tuning 09
"No long-lived credential" IAM role, not user + key 16
"Must not traverse the public internet" PrivateLink, not NAT 16
Two options both in the guide's list The stem's stated constraint breaks the tie 09

Distractor Patterns

Five patterns account for most wrong answers on this paper, and all five recur on the real exam.

1. The true statement that answers a different question. The hardest kind. Q13's option (a) — "only fifty questions affect the result" — is entirely true and simply is not what "compensatory" means. Verifying a claim as true feels like verifying it as the answer. Match the claim to the term being asked about, not to reality.

2. The correct answer to the adjacent question. Q38 offers the two canonical RAG cases as distractors in a fine-tuning question. These are attractive because they are right — about something else. Whenever two techniques are routinely compared, expect each to appear as the other's distractor.

3. The overstated truth. "They cannot process any language other than the one they were trained on." Multilingual capability genuinely varies; "cannot" is false. Check every absolute — always, never, cannot, guarantees, all. On this exam an absolute is wrong far more often than right.

4. The fabricated regulation. "Mandatory disclosure of model weights on request." "Carbon reporting legally required before deployment." No such general rules exist. Domain 4 uses these repeatedly, and "no such rule exists" is a legitimate elimination — you are not required to recognise every regulation, only to notice when one has been invented.

5. The name that sounds adjacent. AgentCore offered as a vector database; Neptune offered where the stem says PostgreSQL. The guard is knowing the actual lists — the vector stores, the governance services — well enough that a plausible-sounding name does not substitute for membership.

Scenario Walkthrough

A media company runs a customer assistant on a managed foundation model. It must answer from the company's own product documentation, which changes several times a week. Legal requires every answer to cite the passage it came from. Security requires that traffic from the application's VPC to the model never traverses the public internet. The team has no ML engineers.

Four requirements, four decisions. Work them separately — the exam bundles them precisely because candidates try to solve the whole scenario with one idea.

"Answer from our own documentation, changing several times a week" plus "cite the passage"RAG, not fine-tuning. Two independent signals point the same way: content that changes frequently would require constant retraining, and per-answer citation requires retrieved passages to cite. Fine-tuning absorbs knowledge into weights, where it cannot be cited and cannot be updated cheaply. (Chapters 09, 11.)

"Managed foundation model" plus "no ML engineers"Bedrock with Knowledge Bases, not SageMaker AI with a self-managed server. Both would work; the stated operational constraint selects the managed path. (Chapter 07.)

"Never traverses the public internet"PrivateLink. Not a NAT gateway, which routes traffic out to the internet, and not a security-group rule, which restricts destinations without changing the path. (Chapter 16.)

Unstated but implied — how do we know it works? The business metric, not the benchmark score. If the assistant deflects support cases, that is the evidence; a leaderboard result is not. (Chapters 07, 12.)

Notice what the walkthrough did not do: it never chose a service because it was the most powerful option. Every decision was made by the stem's stated constraint.

Key Concepts Table

Concept The one-line version Chapter
Compensatory scoring Domains pool into one score; no per-domain bar 01
Unscored trial items 15 of 65, unidentifiable; treat all as scored 01
Multiple response No partial credit; half right scores zero 18
The constraint clause The one phrase in the stem that decides the answer all
Cost ladder In-context → RAG → fine-tuning → pre-training 09
RAG vs fine-tuning Changing content and citation → RAG; behaviour and vocabulary → fine-tuning 09, 11
Config vs CloudTrail State over a period vs who called which API 17
Clarify vs Model Monitor Before the model exists vs after it is deployed 14
Fairness vs safety Harm distributed unevenly vs harmful output generally 13
Capability vs value Benchmarks measure the first; business metrics the second 12

Revision Flashcards

Say the answer aloud before revealing it.

1. The clock is nearly out and you have twelve unanswered questions. What do you do? → Put an answer on all twelve immediately, then refine whatever time remains. An unanswered question scores zero with certainty; a guess between four options returns 25% on average, and 50% once you have eliminated two. Leaving the paper with a blank is the only choice available to you with a guaranteed-zero return, and there is no penalty for a wrong answer.

2. A question says "select two" and you are confident of only one. What now? → Pick a second anyway. Partial credit does not exist on this exam, so a single selection scores exactly what a wrong pair scores — nothing. There is nothing to lose and one option in four to gain. Then count your selections before moving on: an unselected second answer is the most common unforced error on the paper, and it is entirely mechanical.

3. Your practice score is 44 out of 65. Are you ready? → Borderline, and the raw number is the least useful thing you have. Do not re-sit this paper — a second sitting measures recall of these 65 items and nothing else. Take the per-domain totals to the chapters behind each miss, and sort the misses into knowledge gaps and reading errors first, because those two findings have completely different fixes.

4. A question names a service you have never heard of. What does that tell you? → Possibly nothing at all. It may be one of the 15 unscored trial items, or an invented distractor of the kind Domain 4 uses repeatedly. Eliminate whatever you can justify rejecting, answer, flag it, and move on. The real risk is not the question itself — candidates lose marks on the questions that follow a demoralising one far more often than on the demoralising question.

5. Two options are both real services from the correct list. How do you choose? → Re-read the stem for a constraint you have not used yet. When both options are genuinely in scope, the tie-breaker is always something the scenario stated — Neptune really is in the vector-storage list, and it loses to Aurora only because the stem says the team already runs PostgreSQL. Membership of the list never breaks the tie; the constraint does.

6. What is the fastest way to lose marks you have already earned? → Spending disproportionate time on a hard question in a weak domain. Scoring is compensatory and there is no per-domain pass bar, so a mark in your weakest domain is worth exactly what an easy mark elsewhere is worth — and it costs you far more to earn. Bank the cheap marks first; the exam does not care where your marks came from.

Exam-Ready Model Answer

"How should I approach the exam on the day?"

Make one pass through all 65 questions in order, answering every question you can resolve on the first read and flagging everything else — but always recording an answer before moving on, because an unanswered question and a wrong one score the same, and only one of them can turn out lucky. On each question, find the constraint clause in the stem before reading the options, then read the options against it. Where two options survive, choose the one the stated constraint selects rather than the one that sounds most complete.

Because scoring is compensatory and there is no per-domain pass bar, I will not spend disproportionate time on a weak domain: a mark in Domain 4 is worth exactly a mark in Domain 3 and costs me more to earn. On the second pass I revisit flagged questions in flag order, and I finish by confirming that every question carries an answer and that every select-two question carries exactly two.

Chapter Checklist

  • I sat the 65-question paper closed-book, in one sitting, without pausing
  • I recorded an answer for every question before checking any answer
  • I computed per-domain totals, not just the headline score
  • I read the walkthrough for every question — including the ones I got right
  • I identified, for each miss, which of the five distractor patterns caught me
  • I listed the specific chapters my misses point at, by number
  • I can state what compensatory scoring means and what it implies for triage
  • I can name the four question types and the trap in each
  • I know that a select-two question scores zero for a single selection

After the Session

Work the chapters your misses point at — not the whole course, and not this paper again. The walkthroughs name a chapter for every question precisely so that revision can be specific.

If your per-domain totals are weak in Domain 3, start there regardless of the others: it carries 28% of the exam, more than Domains 4 and 5 together, and it is a single connected body of knowledge rather than two.

Confirm the exam's duration and the current guide version on the official exam page before you book. This course was authored against v1.1 (2026-04-30) and the lifecycle was verified 2026-08-07; if you are reading this materially later than that, re-check before relying on the domain weights, because they are what this paper's sampling is built on.

Domain quiz

A team describes their system as "deep learning" in a design review. Which characteristic most reliably confirms that description rather than merely "machine learning"?

A logistics company wants to predict the number of packages a depot will receive tomorrow. What kind of ML problem is this, read from the shape of the required output?

Which scenario is the clearest case for **rejecting** an AI approach outright?

In an ML context, what distinguishes **inference** from **training**?

A retailer holds transaction records with no labels and wants to discover natural customer groupings without deciding the groups in advance. Which approach fits?

A model is described as performing **batch inference** rather than real-time inference. What does this tell you about the workload?

Which of the following is **structured data**?

A team wants a model that improves by taking actions in an environment and receiving a reward signal. Which learning paradigm is this?

Which statement about **agentic AI** is accurate as the term is used in this exam guide?

A scenario states that the business needs to "understand why customers are cancelling, grouping them by behaviour we have not defined in advance." Which output shape does this describe?

Which are genuine characteristics of **computer vision** tasks as scoped by the exam guide? (Select two.)

Choose 2 0 selected

A team must choose between **traditional ML** and a **foundation model** for extracting a fixed set of five fields from a high volume of standardised forms. Which factor most argues for traditional ML here?

What does the **compensatory scoring model** on this exam mean in practice?

What is a **token** in the context of a large language model?

Why do **embeddings** make semantic search possible?

An application must answer questions over a 400-page policy manual. Why is **chunking** the manual necessary before creating embeddings?

A model generates a confident, fluent answer that is factually wrong. What is this called, and what does it indicate?

Which of these is a **diffusion model** best suited to?

A team's Bedrock costs are far higher than forecast. Their prompts each prepend the same 8,000-token policy document before a short user question. Which change most directly reduces cost?

What does **context engineering** refer to in a foundation-model application?

In the **Model Context Protocol**, what problem is being solved?

Which are genuine limitations of foundation models that the exam guide expects you to recognise? (Select two.)

Choose 2 0 selected

A multi-agent system is described in which one agent decomposes a request and delegates sub-tasks to specialist agents. What pattern is this?

Why does **memory management** matter in an agentic system?

An organisation needs a conversational assistant over its own documents, in production, with minimal infrastructure to manage. Which combination is the most direct fit?

What most directly distinguishes **Amazon Bedrock** from **Amazon SageMaker AI** for a team choosing between them?

Which business metric best demonstrates that a deployed GenAI assistant earned its cost?

A team needs the model to follow a fixed output structure every time and to vary its wording as little as possible. Which inference setting most directly serves this?

What is the most accurate description of **Amazon Bedrock AgentCore** in the v1.1 stack?

A team must select a foundation model for an application that processes both scanned images and their accompanying text. Which selection criterion is decisive here?

Raising the **temperature** parameter has what effect on a model's output?

What problem does **Retrieval Augmented Generation** solve that a larger context window alone does not?

An organisation needs to store embeddings for a RAG application and already runs PostgreSQL. Which AWS option most directly supports vector storage there?

Rank these customisation approaches from cheapest to most expensive, as the exam guide frames the cost ladder.

A prompt includes three worked examples before the actual request. What technique is this?

A user submits input designed to override the application's system instructions and reveal its configuration. What is this attack called?

Which evaluation metric is designed for **summarisation** quality?

Which are legitimate reasons to choose **fine-tuning** over RAG? (Select two.)

Choose 2 0 selected

What does **model distillation** produce?

A RAG application returns answers that are fluent but cite passages unrelated to the question. Where should the team look first?

Why is **human-in-the-loop** review used when evaluating a generative model?

An application needs the same lengthy system instruction on every request, and the team wants to manage its versions centrally rather than embedding it in application code. Which capability serves this?

What is **continuous pre-training**, as distinct from fine-tuning?

A team reports a strong benchmark score and asks to declare the project a success. What is the correct objection?

Which factor most directly argues for a **smaller** foundation model in a production application?

In **in-context learning**, where does the task-specific knowledge live?

What is the primary role of an **AI agent** in a business application?

A hiring model performs well overall but rejects qualified candidates from one region at a much higher rate. Which responsible-AI property is failing?

Which AWS capability most directly detects statistical bias in a training dataset before a model is built?

A model achieves 99% training accuracy and 61% test accuracy. What does this indicate?

Why does **dataset inclusivity** matter beyond fairness as an abstract principle?

What is the purpose of an **Amazon SageMaker Model Card**?

Which are genuine legal risks specific to generative AI named in the exam guide? (Select two.)

Choose 2 0 selected

A team argues that a more transparent model would let attackers probe its decision boundary. What does this describe?

Which best describes **human-centered design** for explainable AI?

Why is **environmental impact** treated as a genuine model-selection criterion rather than a footnote?

An application must call Bedrock without any long-lived credential stored on the compute instance. What is the correct mechanism?

Which service most directly discovers sensitive personal data sitting in an S3 bucket used as a training corpus?

Traffic between a VPC and Bedrock must not traverse the public internet. Which mechanism achieves this?

Under the **shared responsibility model**, which is the customer's responsibility when using a managed foundation model service?

Which technique most directly reduces hallucination in a deployed assistant?

Which practices belong to **secure data engineering** for an AI workload? (Select two.)

Choose 2 0 selected

A regulator asks which identity disabled encryption on an inference endpoint, and when. Which service answers this?

What is the purpose of the **Generative AI Security Scoping Matrix**?

An organisation must prove its data residency commitment for a model trained on EU customer data. Which governance control is most directly relevant?

← Back to AIF-C01