AWS interview questions in 2026 test whether you can design, secure and pay for real systems, not whether you can recite service names. Expect core questions on EC2, S3, IAM and VPC, serverless questions on Lambda, API Gateway and DynamoDB, and at least one scenario such as "design a serverless image pipeline" or "cut this bill by 40%." The candidates who pass explain trade-offs, failure modes and cost in the same answer.
This guide gives you model answers for each area, two fully worked scenarios with reference diagrams, and decision tables you can reuse in any loop.
Key Takeaways
- Every strong AWS answer follows the same shape: requirements, request path, service choice with the rejected alternative, failure modes, then cost.
- IAM and VPC questions eliminate more candidates than Lambda trivia. Know roles vs users, policy evaluation order, and why a private subnet still needs a NAT gateway or endpoint.
- Lambda has hard limits that drive design: a 15-minute timeout, up to 10,240 MB of memory, and a 6 MB synchronous payload. Large files go through S3, not through the function.
- Multi-AZ is the default bar for production. Multi-region is a deliberate, expensive choice you should justify with a recovery target.
- Cost questions reward a method: find the top line items, rightsize, commit with Savings Plans, use Spot where interruption is safe, and kill NAT and data transfer waste.
What Do AWS Interviews Look Like in 2026?
AWS questions show up in three kinds of loops. Cloud and DevOps roles get a dedicated AWS round. Backend roles fold AWS into a system design interview. Solutions architect roles get several scenario rounds plus customer-style conversations.
The difficulty scales with level. Juniors explain services. Mid-level engineers compare options. Senior engineers and anyone answering "AWS interview questions for experienced" prompts must design under constraints and defend cost. If you are interviewing at Amazon itself, expect the technical answers to be scored alongside the Amazon Leadership Principles, especially Frugality and Dive Deep.
| Role | Typical AWS depth | What gets scored |
|---|---|---|
| Backend engineer | Medium | Data store choice, queues, API design on AWS |
| DevOps / platform | High | IAM, VPC, IaC, CI/CD, observability |
| Site reliability engineer | High | Multi-AZ, failover, alarms, incident stories |
| Solutions architect | Very high | Scenarios, cost, migration plans, trade-offs |
Core Services: EC2, S3, IAM and VPC Questions
These questions open most AWS rounds. Answer in two or three sentences, then add one production detail to show you have used the service.
EC2 and storage
What is the difference between EBS, instance store and EFS? EBS is network-attached block storage tied to one Availability Zone that persists after the instance stops. Instance store is physically attached disk that is fast but lost when the instance stops or fails. EFS is a managed NFS file system that many instances across AZs can mount at once.
How does an Auto Scaling group decide to add instances? It follows a scaling policy. Target tracking is the common default: you set a goal like 50% average CPU or a request count per target, and the group adds or removes instances to hold it. Mention health checks from the load balancer so unhealthy instances get replaced, not just counted.
S3
Is S3 eventually consistent? No. Since December 2020, S3 provides strong read-after-write consistency for all PUT and DELETE operations at no extra cost. Many older prep guides still say otherwise, and interviewers notice.
How do you choose an S3 storage class? Match it to access frequency. Standard for hot data, Intelligent-Tiering when access patterns are unknown, Standard-IA for infrequent access with fast retrieval, and the Glacier classes for archives. Lifecycle rules move objects automatically.
IAM
IAM is where experienced candidates separate themselves. The questions sound basic but the follow-ups go deep.
Role vs user? A user has long-term credentials and represents a person or legacy system. A role has no long-term credentials; any trusted principal (an EC2 instance, a Lambda function, a federated human) assumes it and gets temporary credentials from STS. Production workloads should use roles.
How is a request evaluated? Everything starts as an implicit deny. An explicit deny anywhere wins. Otherwise the request needs an allow from the identity or resource policy, and it must also pass any boundaries in play: Service Control Policies, permission boundaries and session policies.
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:GetObject"],
"Resource": "arn:aws:s3:::uploads-bucket/raw/*"
}
]
}
Use a snippet like this to show least privilege: one action, one prefix, no wildcards on the bucket.
AWS VPC Interview Questions: Networking and Security Scenarios
VPC questions are where most candidates lose points, because they need a mental model of the packet path. A VPC is a logically isolated network in one region, split into subnets that each live in a single AZ.
What makes a subnet public? Its route table has a route to an internet gateway. Nothing else. A private subnet has no such route, so instances there reach the internet through a NAT gateway in a public subnet.
Security group vs NACL? Security groups are stateful, attach to network interfaces and only allow. NACLs are stateless, apply at the subnet edge, support deny rules and need ephemeral return ports opened.
Scenario: a Lambda in a private subnet times out calling S3. Why? The function has no path out. Private subnets have no internet route, so either add a NAT gateway or, better, add an S3 gateway endpoint to the route table. Gateway endpoints for S3 and DynamoDB carry no hourly or data processing charge, while NAT gateways bill per hour and per GB processed.
| Need | Use | Why |
|---|---|---|
| Private subnet to S3 or DynamoDB | Gateway VPC endpoint | No charge, traffic stays on AWS network |
| Private subnet to other AWS APIs (SQS, Secrets Manager) | Interface endpoint (PrivateLink) | Private IP in your subnet, per-hour plus per-GB cost |
| Private subnet to the public internet | NAT gateway | Outbound only, managed, per-AZ |
| VPC to VPC, a few pairs | VPC peering | Simple, non-transitive |
| Many VPCs and on-prem | Transit Gateway | Hub-and-spoke, transitive routing |
| Expose one service to other accounts | PrivateLink endpoint service | No full network connectivity needed |
A follow-up that catches people: "Why deploy one NAT gateway per AZ?" Because a NAT gateway lives in one AZ. If that AZ fails, private subnets in other AZs routing through it lose outbound access, and cross-AZ traffic adds data transfer charges.
AWS Lambda Interview Questions: Serverless with API Gateway and DynamoDB
Lambda is a managed compute service that runs your function in response to events and bills per request and per millisecond of duration. Interviewers want to see that you know where it breaks.
Know the limits that shape designs, per the Lambda quotas documentation: a 15-minute maximum timeout, memory from 128 MB to 10,240 MB (CPU scales with memory), up to 10 GB of ephemeral /tmp storage, and a 6 MB payload for synchronous invocations. API Gateway REST APIs default to a 29-second integration timeout. AWS has since let regional and private REST APIs request a higher limit, but long jobs still belong behind a queue.
How do you reduce cold starts? Keep the deployment package small, initialize clients outside the handler so warm invocations reuse them, prefer lighter runtimes for latency-critical paths, and use provisioned concurrency where p99 latency matters. SnapStart reduces startup time for supported runtimes such as Java.
How do you make a Lambda consumer idempotent? Assume at-least-once delivery from SQS, SNS and EventBridge. Store a processed-event key in DynamoDB with a conditional write, and skip duplicates.
import boto3
from botocore.exceptions import ClientError
table = boto3.resource("dynamodb").Table("processed_events")
def handler(event, context):
for record in event["Records"]:
key = record["messageId"]
try:
table.put_item(
Item={"pk": key},
ConditionExpression="attribute_not_exists(pk)",
)
except ClientError as e:
if e.response["Error"]["Code"] == "ConditionalCheckFailedException":
continue
raise
process(record)
How do you model data in DynamoDB? Start from access patterns, not entities. Pick a partition key with high cardinality so traffic spreads evenly, use the sort key for range queries, and add global secondary indexes for alternate lookups. Items cap at 400 KB, so large blobs go to S3 with a pointer in the item. Hot partitions are the classic failure; the fix is a better key or write sharding. If the follow-up drifts toward caching hot reads, our distributed cache design guide covers DAX-style and Redis patterns.
| Workload | First choice | Choose instead when |
|---|---|---|
| Spiky, event-driven, short tasks | Lambda | Steady high RPS makes per-request cost higher than containers |
| Long-running HTTP service | ECS on Fargate | Team already runs Kubernetes, then EKS (Kubernetes questions here) |
| Key-value at any scale, known access patterns | DynamoDB | Ad hoc joins and reporting needed, then Aurora or RDS |
| Relational, transactional | Aurora / RDS | Global single-digit-ms key lookups, then DynamoDB |
| Decoupling producer and consumer | SQS | Fan-out to many consumers, then SNS or EventBridge |
| Ordered event stream with replay | Kinesis or MSK | Simple work queue, then SQS |
Scenario 1: Design a Serverless Image Pipeline
This is one of the most common AWS scenario-based interview questions. Prompt: users upload photos from a mobile app; generate three thumbnail sizes, store metadata, and handle traffic spikes of 50x normal.
Start with clarifying questions: maximum file size, latency target for thumbnails, retention, and whether uploads are public or authenticated. Then draw the path.
Mobile app
| 1. POST /uploads (Cognito JWT)
v
API Gateway --> Lambda "presign" --> returns S3 presigned PUT URL
|
| 2. PUT image directly to S3 (bypasses API Gateway payload limits)
v
S3 bucket: raw/ --(ObjectCreated event)--> SQS queue --> DLQ after N failures
|
v
Lambda "resize" (batch, reserved concurrency)
|
+-----------------------+----------------------+
v v
S3 bucket: thumbs/ DynamoDB: image metadata
|
v
CloudFront (OAC) --> clients
Explain each choice and the alternative you rejected:
- Presigned URLs instead of uploading through API Gateway. API Gateway and Lambda payload limits make large images fail, and streaming bytes through a function wastes compute.
- SQS between S3 and the resize function. The queue absorbs the 50x spike, gives retries with a dead-letter queue, and lets you cap concurrency so you do not exhaust account concurrency or a downstream service.
- Reserved concurrency on the resize function. Protects other functions in the account during a spike.
- Separate prefixes or buckets for raw and thumbnails. Writing thumbnails back to the trigger prefix creates an infinite loop, a classic trap interviewers wait for you to mention.
- CloudFront with origin access control. The bucket stays private and reads are cached at the edge.
Close with failure modes: poison images go to the DLQ with an alarm, idempotent writes use the S3 key as the DynamoDB key, and lifecycle rules move raw originals to a cheaper class after 30 days. This kind of event pipeline looks a lot like the fan-out in a notification system design, so the same reasoning transfers.
Scenario rounds punish blank moments. TechScreen listens to the prompt and surfaces a structured architecture outline on your screen in real time, invisible during Zoom, Meet and Teams screen shares. Try it on a mock AWS round with 3 free tokens, no credit card.
Reliability and Multi-AZ Design Questions
Reliability questions test whether you know which failures AWS handles for you and which you must design for. An Availability Zone is one or more discrete data centers with independent power and networking inside a region.
How do you make a web tier highly available? Put an Application Load Balancer in front of an Auto Scaling group spread across at least two, ideally three, AZs. Keep instances stateless, store sessions in ElastiCache or DynamoDB, and let health checks replace failed hosts.
RDS Multi-AZ vs read replicas? A classic Multi-AZ instance deployment keeps a synchronous standby in another AZ for automatic failover; it improves availability, not read throughput (Multi-AZ DB clusters add readable standbys). Read replicas use asynchronous replication to scale reads and can be promoted manually. Aurora stores six copies of data across three AZs at the storage layer.
How do you pick a disaster recovery strategy? Start from RTO (how long you can be down) and RPO (how much data you can lose), then pick the cheapest pattern that meets them.
| Strategy | Typical RTO | Typical RPO | Relative cost |
|---|---|---|---|
| Backup and restore | Hours | Hours | Lowest |
| Pilot light | Tens of minutes | Minutes | Low |
| Warm standby | Minutes | Seconds to minutes | Medium |
| Multi-site active/active | Near zero | Near zero | Highest |
The AWS Well-Architected Framework reliability pillar uses these same patterns, so naming them signals you know the shared vocabulary. Reliability also covers protecting downstream services; be ready to sketch throttling with API Gateway usage plans or a token bucket, as in our rate limiter design walkthrough. For on-call and SLO depth, see the site reliability engineer interview guide.
Cost Optimization Questions and the "Cut This Bill by 40%" Scenario
Cost questions reward a method more than a list. Prompt: "Our AWS bill is $100k a month. Cut it by 40% without hurting reliability." Use this sample breakdown to reason aloud, and say you would confirm real numbers in Cost Explorer first.
| Line item (illustrative) | Monthly | Action | Plausible saving |
|---|---|---|---|
| EC2 on-demand, steady web fleet | $40k | Rightsize from metrics, move to Graviton where compatible, then cover the baseline with a Compute Savings Plan | $14k-18k |
| NAT gateway data processing | $12k | Gateway endpoints for S3 and DynamoDB, keep traffic in-AZ | $6k-9k |
| RDS | $15k | Rightsize, delete unused replicas, reserved instances for the steady primary | $3k-5k |
| S3 | $10k | Lifecycle to IA and Glacier, Intelligent-Tiering for unknown access, expire old versions | $3k-5k |
| Batch / CI workers | $8k | Move to Spot with interruption handling | $4k-6k |
| Idle and orphaned resources | $5k | Unattached EBS volumes, old snapshots, unused Elastic IPs and load balancers | $4k-5k |
| Data transfer and logs | $10k | CloudFront for egress, shorter CloudWatch log retention | $2k-4k |
The total lands around $36k-52k, which covers the 40% target with a margin. Make the order explicit: quick wins first (idle resources, endpoints, lifecycle rules), then rightsizing, then commitments. Committing to Savings Plans before rightsizing locks in waste.
Three details earn extra credit. AWS charges for public IPv4 addresses (since February 2024, about $0.005 per hour each), which adds up across large fleets. Cross-AZ traffic is billed in both directions, so chatty services split across AZs cost more than you expect. And Spot savings only count if the workload checkpoints or retries cleanly; say how you would handle the two-minute interruption notice.
How to Structure Any AWS Scenario Answer
Use the same five-step frame for every scenario so you never stall.
- Requirements. Users, traffic, data size, latency, compliance, budget. Ask, do not assume.
- Request path. Draw client to edge to compute to storage. Name every hop.
- Choices with rejected alternatives. "SQS, not Kinesis, because we need a work queue without ordering or replay."
- Failure modes. AZ loss, throttling, poison messages, retries without idempotency, IAM misconfiguration.
- Cost and operations. Biggest cost driver, how it scales, and which alarms you would set.
Practice this frame on classic prompts like a URL shortener on AWS, then layer in AWS-specific services. If you are prepping for a backend-heavy loop, pair this with the backend engineer interview guide. And if the role is at Amazon, read the Amazon technical interview process breakdown, since the bar raiser will probe the same trade-offs.
Final Prep Checklist
- Build one small project end to end: presigned upload, SQS, Lambda, DynamoDB, CloudFront. Break it on purpose and fix it.
- Write two IAM policies from memory, one identity policy and one bucket policy with a condition.
- Draw a three-AZ VPC with public and private subnets, NAT per AZ and gateway endpoints.
- Prepare one real cost story with a number, and one incident story with a timeline.
- Rehearse the five-step scenario frame aloud on three prompts until it takes under two minutes to reach a first diagram.
Want a second brain in your next AWS interview? TechScreen runs quietly on your desktop, stays invisible on screen shares in Zoom, Google Meet, Teams, HackerRank and CoderPad, and helps you structure VPC, Lambda and cost answers in real time. Start with 3 free tokens, no credit card required.
Frequently Asked Questions
What AWS topics come up most in interviews?
Most AWS interviews cover five areas: compute and storage basics (EC2, S3, EBS), identity and access (IAM roles and policies), networking (VPC, subnets, security groups, NAT, endpoints), serverless (Lambda, API Gateway, DynamoDB, SQS), and reliability plus cost. Experienced candidates also get scenario questions where they design a system or reduce a bill, and interviewers grade the trade-offs they name more than the services they list.
How do I answer AWS scenario-based interview questions?
Start by restating the requirements and asking about scale, latency, budget and compliance. Then sketch the request path from client to storage, name each service and say why you picked it over the obvious alternative. Finish with failure modes and cost. A clear structure of requirements, diagram, choices, failures and cost works for almost every AWS scenario question and keeps you from listing services without reasons.
Do I need an AWS certification to pass an AWS interview?
No. A certification such as Solutions Architect Associate helps get past resume filters and gives you shared vocabulary, but interviewers rarely ask about it. They want to hear how you operated real workloads: an outage you debugged, a permission boundary you designed, or a bill you reduced. Hands-on stories beat certificates, so build and break a small project if you lack production experience.
What is the difference between a security group and a network ACL?
A security group is a stateful firewall attached to an elastic network interface, so return traffic is allowed automatically and it only supports allow rules. A network ACL is a stateless filter at the subnet boundary, evaluated in rule-number order, with both allow and deny rules, so you must open ephemeral return ports explicitly. In practice, security groups do most of the work and NACLs act as a coarse subnet-level guard.
When should I choose Lambda over containers on ECS or EKS?
Choose Lambda for event-driven or spiky workloads with short tasks under the 15-minute limit, where paying per invocation beats paying for idle capacity. Choose ECS or EKS for long-running services, steady high traffic where per-request pricing gets expensive, workloads needing special hardware, or teams that need fine control over the runtime. Many production systems mix both.
What are the most common AWS cost optimization answers interviewers expect?
Interviewers expect you to start with visibility: tags, Cost Explorer and finding the top few line items. Then rightsize instances, move steady usage to Savings Plans, use Spot for fault-tolerant jobs, add S3 lifecycle rules, replace NAT gateway traffic to S3 and DynamoDB with free gateway endpoints, and delete idle resources such as unattached volumes and old snapshots. Always tie each step to a measured saving.
Ready to use AI assistance in your next interview?
TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.
Ace your next interview →