StudyToCert

All certifications / Developer Associate / Lessons

AWS Certified Developer – Associate DVA-C02 · Domain 1: Development with AWS Services

Resilient code: retries with exponential backoff and jitter, idempotency, timeouts, handling partial failures and dead-letter queues

▶ Watch the overview video

Last reviewed September 25, 2026 · Leer en español

Distributed systems fail in small ways all the time: a request is throttled, a network call times out, a downstream service briefly returns an error. Resilient code expects this. The DVA-C02 exam tests whether you know the standard techniques and which failures they fix.

Retries are the first tool, but only for transient errors such as throttling (ThrottlingException, HTTP 429), HTTP 5xx server errors and network timeouts. Retrying a validation error or an AccessDenied just repeats the failure. Retrying immediately in a tight loop is harmful, because every client hammers the struggling service at once. Exponential backoff waits longer after each failed attempt, for example 100 ms, 200 ms, 400 ms, 800 ms, up to a cap and a maximum number of attempts. Jitter adds randomness to each wait so that thousands of clients that failed at the same moment do not all retry at the same moment. The AWS SDKs already implement retries with backoff and jitter for AWS API calls; you configure the retry mode and maximum attempts rather than writing it yourself, but you add it yourself for calls to your own or third-party services.

Idempotency makes retries safe. An operation is idempotent if performing it twice has the same effect as performing it once. Because retries, at-least-once message delivery in SQS standard queues and asynchronous Lambda retries can all deliver the same request more than once, your code should detect duplicates. Common techniques are an idempotency key supplied by the client (for example an order ID), a DynamoDB conditional write such as attribute_not_exists(orderId) that fails if the item was already processed, and natural idempotency (setting a value rather than incrementing it).

Timeouts stop one slow dependency from consuming all your resources. Set explicit connect and read timeouts on HTTP and SDK clients that are shorter than your function or request timeout, so your code can log, retry or fail gracefully instead of being killed mid-operation. A Lambda function behind API Gateway, for instance, should give up on a slow downstream call well before API Gateway's own integration timeout.

Partial failure happens when a batch contains good and bad items. If one record in a batch of ten fails and you throw an error, the whole batch is retried, including the nine that succeeded. Batch APIs such as DynamoDB BatchWriteItem return UnprocessedItems and SQS SendMessageBatch returns per-entry failures; your code must retry only those. For Lambda reading from SQS or streams, partial batch responses let you report just the failed items.

A dead-letter queue (DLQ) is where messages go after they have failed a set number of times, so a poison message (one that can never be processed) stops blocking the queue and wasting compute. SQS uses a redrive policy with maxReceiveCount; asynchronous Lambda invocations can send failed events to an SQS queue or SNS topic DLQ, or to an on-failure destination. A DLQ is only useful if you alarm on its depth and investigate what lands there.

Key terms

Exponential backoff
A retry strategy where the wait between attempts grows multiplicatively, reducing pressure on a struggling service.
Jitter
Random variation added to retry delays so many clients do not retry in synchronized waves.
Idempotency
The property that repeating an operation produces the same result as doing it once, making retries and duplicate deliveries safe.
Poison message
A message that fails processing every time it is received and would loop forever without a dead-letter queue.
Dead-letter queue (DLQ)
A queue that receives messages or events that failed processing after a configured number of attempts, for later inspection.
Real-world example

A payment Lambda function processes SQS messages. It records each payment ID in DynamoDB with a conditional put using attribute_not_exists, so a message delivered twice is charged once. Its SQS queue has a redrive policy with maxReceiveCount of 5 and a DLQ, and a CloudWatch alarm fires when the DLQ holds any messages.

Exam tip: If a question mentions throttling errors such as ProvisionedThroughputExceededException or ThrottlingException, the answer is almost always retries with exponential backoff (and jitter), not simply adding more retries or raising timeouts.

Check yourself

Why add jitter to exponential backoff?

Without randomness, clients that failed together retry together, recreating the load spike; jitter spreads retries out over time.

Which errors should not be retried?

Client errors that will fail again unchanged, such as validation errors, malformed requests or AccessDenied; retry only transient errors like throttling, 5xx and timeouts.

How does a DLQ help with a poison message in SQS?

After the message's receive count exceeds maxReceiveCount, SQS moves it to the DLQ, so it stops being retried and blocking processing, and you can inspect it.

Study Developer Associate for free
A week-by-week plan with every lesson, quizzes, checkpoint tests, a practice exam and hands-on labs.
Open the Developer Associate study plan