← All articles
13 min read

Backend Engineer Interview Questions and Prep Guide (2026)

Backend loops test far more than LeetCode. This guide maps what each round checks at each level and walks through the API design round most prep guides skip.

Backend engineer interview questions in 2026 cover four areas: coding (data structures and algorithms), API design, data storage (indexing, transactions, sharding), and distributed system design with caching, queues, and failure handling. To pass, you need to solve medium-level coding problems cleanly, design an API with sensible pagination and error handling, defend your database choices under concurrent load, and scale your design to the level you are interviewing for.

This guide breaks down the typical backend loop, gives you a skills matrix of what each level is expected to show, and walks through a full API design round, which most prep guides skip entirely.

Key Takeaways

  • A typical backend loop is a recruiter call, one technical screen, and a 4 to 5 round onsite: 1 to 2 coding, 1 system design, often 1 API or schema design, and 1 behavioral.
  • The bar changes by level more than the questions do. A junior candidate is graded on correctness; a senior candidate on tradeoffs, failure modes, and operability.
  • In API design rounds, cursor pagination, idempotency keys, and a clear versioning policy separate strong answers from average ones.
  • Database depth is the most common gap: know composite index ordering, isolation levels and their anomalies, and how to pick a shard key.
  • Every design answer should say what happens when a dependency is slow, a message is delivered twice, or a cache goes cold.
  • Six weeks of structured prep at about ten hours per week is enough for most working engineers.

What Does a Typical Backend Engineer Interview Loop Look Like?

A backend interview loop is a sequence of rounds that tests coding, data modeling, distributed design, and collaboration, usually over three to six weeks. The structure at most mid-size and large companies looks like this:

StageLengthWhat it tests
Recruiter screen20-30 minBackground, level fit, stack, compensation expectations
Technical screen45-60 minOne or two coding problems, sometimes a short design discussion
Coding rounds (1-2)45-60 min eachAlgorithms, clean code, testing, complexity
System design45-60 minScaling a service end to end
API or schema design45-60 minEndpoints, data model, pagination, errors, evolution
Behavioral45 minOwnership, conflict, incidents, impact

Many companies fold API design into system design instead of running a separate round. Companies with practical loops swap algorithm rounds for work-like tasks: the Stripe interview process is the best-known example, with an integration round and a debugging round instead of puzzles.

Backend coding rounds draw from the same problem pool as everyone else. The coding interview patterns cheat sheet covers the fifteen or so patterns that handle most of them.

The Backend Skills Matrix: What Each Level Is Graded On

Interviewers use similar prompts at every level and adjust the bar. This matrix reflects the expectations commonly described in public interview rubrics and engineering ladders; exact wording differs by company.

SkillJunior / New gradMid-levelSeniorStaff
CodingCorrect, tested solution to a medium problemClean, idiomatic code, edge cases unpromptedFast and clean, discusses production concernsSame as senior; rarely the deciding round
API designReasonable CRUD endpoints and status codesPagination, validation, error formatIdempotency, versioning, backward compatibilityAPI governance across teams, deprecation strategy
DatabasesSQL joins, basic indexesComposite indexes, transactions, N+1 queriesIsolation levels, locking, replication lag, shard keysMigrations at scale, multi-region data, cost
Caching and queuesKnows what a cache is forCache-aside, TTLs, basic queue usageInvalidation, stampedes, delivery semantics, DLQsChoosing org-wide platforms and their limits
Concurrency and reliabilityRace conditions in principleLocks, timeouts, retriesBackoff with jitter, circuit breakers, idempotent consumersFailure domains, SLOs, graceful degradation strategy
System designUsually not tested, or a light versionOne service plus database plus cacheFull distributed design, drives the conversationAmbiguous problems, multi-year evolution, org impact
BehavioralLearning, teamworkOwnership of featuresLeading projects, incidents, mentoringInfluence without authority, technical strategy

The practical read: a mid-level candidate who answers like a senior on databases and reliability often gets up-leveled. A senior candidate who answers system design like a mid-level (one box per component, no failure discussion) often gets down-leveled or rejected.

How Does the API Design Round Work?

The API design round asks you to define the interface for a service: resources, endpoints or RPCs, request and response shapes, errors, pagination, and how the API changes over time. It is usually 45 to 60 minutes, and the prompt is concrete, such as "design the API for a task management app" or "design an API for placing and tracking orders."

Use this five-step structure:

  1. Clarify clients and use cases (5 min). Who calls this: a mobile app, a browser, partners, internal services? What are the top three operations? What are the read and write volumes?
  2. Pick the protocol (2 min). REST, gRPC, or GraphQL, with one sentence of reasoning.
  3. Model resources (10 min). Nouns, their fields, IDs, and relationships.
  4. Define operations (15 min). Endpoints, payloads, status codes, pagination, filtering.
  5. Cover cross-cutting concerns (15 min). Errors, idempotency, auth, rate limits, versioning.

REST vs gRPC vs GraphQL

FactorREST (JSON over HTTP)gRPC (Protobuf over HTTP/2)GraphQL
Best forPublic and partner APIs, browsersInternal service-to-service callsClient-driven reads across many entities
ContractOpenAPI (optional).proto files, generated clientsTyped schema
PayloadText, human-readableBinary, compactJSON
StreamingLimited (SSE, WebSockets separately)Built-in client, server, and bidirectionalSubscriptions
Browser supportNativeNeeds gRPC-Web or a proxyNative
CachingHTTP caching works wellNo HTTP caching by defaultHarder; usually one POST endpoint
Debuggingcurl and a browserNeeds toolingNeeds tooling

A solid default answer: REST for the public API, gRPC between internal services. Mention GraphQL only if the prompt has many client types fetching nested data.

Worked example: an orders API

A strong REST answer to "design an API to create and list orders":

POST /v1/orders
Idempotency-Key: 7f3c2a90-1b4e-4c8a-9d2f-55e1a0c3b6d1
Content-Type: application/json

{
  "customer_id": "cus_123",
  "items": [{ "sku": "SKU-42", "quantity": 2 }],
  "currency": "usd"
}

HTTP/1.1 201 Created
Location: /v1/orders/ord_789

{
  "id": "ord_789",
  "status": "pending",
  "total_amount": 4998,
  "currency": "usd",
  "created_at": "2026-10-01T14:03:00Z"
}
GET /v1/orders?customer_id=cus_123&status=paid&limit=50&cursor=eyJpZCI6Im9yZF83ODkifQ

HTTP/1.1 200 OK

{
  "data": [ { "id": "ord_788", "status": "paid" } ],
  "next_cursor": "eyJpZCI6Im9yZF83NTAifQ",
  "has_more": true
}

Points to say out loud while writing this:

  • Money as integers in the smallest unit (cents), never floats.
  • 201 with a Location header on create, 200 on reads, 404 for missing resources, 409 for conflicts, 422 or 400 for validation errors, 429 for rate limits.
  • Idempotency key on POST, so a client retry after a timeout does not create two orders. Store the key with the response and return the stored response on a repeat. Stripe's idempotent requests documentation is the reference design most interviewers know.
  • A consistent error envelope, such as a machine-readable code, a human message, and the offending field.

Pagination: offset vs cursor

Many candidates lose points by defaulting to ?page=3. With offset pagination, the database scans and discards every skipped row, and results shift when new rows arrive. Cursor (keyset) pagination encodes the last seen sort key and queries from there:

SELECT id, status, created_at
FROM orders
WHERE customer_id = $1
  AND (created_at, id) < ($2, $3)
ORDER BY created_at DESC, id DESC
LIMIT 51;

Fetch limit + 1 rows to know whether has_more is true. The id tiebreaker matters because created_at is not unique. Keep the cursor opaque so its internals can change. Google's AIP-158 pagination guidance uses the same opaque page token approach. Offset is still fine for small admin tables where users jump to page N.

Versioning and evolution

State a versioning policy instead of just writing /v1/. A good answer covers three things:

  • Where the version lives: URL path (/v1/) is the most common and easiest to route; header-based or date-based versions are alternatives.
  • What counts as breaking: removing or renaming a field, changing a type, or adding a required parameter. Adding optional fields or new endpoints is not breaking.
  • How you retire versions: deprecation notices, usage metrics per version, and a sunset date.

For gRPC, mention the Protobuf rules: never reuse or renumber field tags, reserve removed field numbers, and add fields rather than changing them.

Database Interview Questions: Indexing, Transactions, Sharding

Database questions are where backend loops separate candidates most sharply. Expect them inside system design and as standalone follow-ups. If your SQL is rusty, work through SQL interview questions before going deeper here.

Indexing

An index is a separate data structure, usually a B-tree, that lets the database find rows without scanning the table. The questions that come up most:

  • Composite index order. An index on (customer_id, created_at) serves WHERE customer_id = ? ORDER BY created_at, but not a filter on created_at alone. Put equality columns first, then range or sort columns.
  • Covering indexes. If the index contains every column a query needs, the database can skip reading the table.
  • Write cost. Each index slows writes. Do not index everything.
  • Diagnosis. Use EXPLAIN to confirm the plan and spot full scans.
  • N+1 queries. One query for a list plus one per item. Fix it with a join or a batched IN query.

Transactions and isolation

A transaction groups operations so they all succeed or all fail (ACID: atomicity, consistency, isolation, durability). Interviewers then probe isolation levels:

Isolation levelPreventsStill allows
Read CommittedDirty readsNon-repeatable reads, lost updates, write skew
Repeatable Read / SnapshotDirty and non-repeatable readsWrite skew (in snapshot implementations)
SerializableAll of the above, including write skewNo anomalies, at the cost of retries or blocking

Know the defaults: PostgreSQL uses Read Committed and MySQL's InnoDB uses Repeatable Read. The PostgreSQL transaction isolation docs explain exactly which anomalies each level permits in Postgres.

The classic follow-up is "two users buy the last item at once." Strong answers offer a conditional update (UPDATE inventory SET qty = qty - 1 WHERE sku = ? AND qty > 0, then check rows affected) or optimistic locking with a version column. SELECT ... FOR UPDATE also works but holds locks longer.

Replication and sharding

Replication copies data to other nodes for read scale and failover. The trap is replication lag: a user writes, then reads a stale replica. Fix it by reading recent writes from the primary.

Sharding splits data across nodes by a shard key. A good shard key spreads load evenly and keeps common queries on one shard. user_id usually works for user-centric apps; created_at creates a hot shard because all new writes land in one place. Mention consistent hashing for adding nodes with minimal data movement, and acknowledge the cost: cross-shard joins and transactions become hard, so you design to avoid them.

Caching and Queues

Caching and messaging come up in nearly every backend system design, and interviewers check whether you know how they fail, not just how they help.

Caching

Cache-aside (the application reads the cache, falls back to the database, then populates the cache) is the default pattern to describe. Then cover the failure modes:

  • Invalidation: delete the cache entry on write instead of updating it, to avoid racing writers leaving stale data.
  • Stampede: when a hot key expires, thousands of requests hit the database at once. Use request coalescing, a short lock, or early probabilistic refresh.
  • TTL jitter: randomize expirations so keys do not expire together.

The distributed cache design walkthrough covers eviction, partitioning, and replication if the interviewer asks you to build the cache itself.

Queues and streams

A message queue decouples producers from consumers so slow work happens asynchronously. Know the difference between a work queue (such as SQS or RabbitMQ, where each message is processed by one consumer) and a log-based stream (such as Kafka, where consumers track offsets and messages can be replayed). Kafka guarantees ordering only within a partition, so partition by the key whose order matters, such as order_id.

Delivery semantics are the most common follow-up. Most systems give at-least-once delivery, which means duplicates will happen. The answer is idempotent consumers: record processed message IDs or make the write itself idempotent (an upsert keyed on the event ID). Add a dead-letter queue for messages that keep failing, and alert on its depth.

To publish events reliably alongside a database write, describe the transactional outbox: write the event to an outbox table in the same transaction, then publish it from a separate process. This avoids the database committing while the publish fails.

Backend rounds stack SQL, isolation levels, and distributed design follow-ups faster than anyone can recall every detail on the spot. TechScreen is an invisible real-time AI interview assistant that stays hidden during screen share on Zoom, Google Meet, Teams, and CoderPad, and helps you structure answers as the questions come. Start with 3 free tokens, no credit card required.

Get started free →

Concurrency and Reliability Questions

Backend interviewers want to hear that you expect things to break. The common topics:

  • Race conditions. Check-then-act bugs, lost updates, and double-spend. Fix with atomic operations, database constraints, or locks. Unique constraints are an underrated answer: let the database reject the duplicate.
  • Timeouts. Every network call needs one. Without it, a slow dependency ties up threads until the whole service stalls.
  • Retries with exponential backoff and jitter. Retry only idempotent operations, cap the attempts, and add randomness so clients do not retry in sync.
  • Circuit breakers. Stop calling a failing dependency for a cooldown period and return a fallback, so failures do not cascade.
  • Rate limiting. Token bucket or sliding window, per user or per API key. The rate limiter design guide walks through the algorithms and the distributed version.
  • Distributed locks. A leased lock can expire while the holder is paused; fencing tokens solve this.

Coding rounds for some backend roles include a concurrency problem, such as implementing a thread-safe bounded queue, a rate limiter, or a worker pool. Practice those in your interview language with the concurrency and multithreading interview questions guide.

Backend System Design by Level

Backend system design is the round that most often sets your level. The same prompt, such as "design a payment service" or "design a notification system," is graded differently depending on the role.

LevelWhat a passing answer includesCommon failure
Junior / new gradAPI, data model, one service, one database, a cacheUnable to estimate scale or explain data flow
Mid-levelAbove plus capacity estimate, indexes, caching strategy, async processingNo discussion of failures or bottlenecks
SeniorDrives the session, picks storage with justification, covers consistency, idempotency, failure modes, monitoringGeneric boxes with no depth on the hardest component
StaffFrames the problem, explores alternatives, multi-region, migrations, cost, team boundariesGoes deep too early without aligning on scope

Prompts that come up often for backend roles, with full walkthroughs:

For the general framework (requirements, estimates, high-level design, deep dives, wrap-up), use the system design interview guide.

A senior habit worth copying: after the high-level diagram, name the hardest part (exactly-once effects for payments, fan-out for a feed) and spend most of the remaining time there.

A 6-Week Backend Interview Prep Plan

This plan assumes about ten hours per week while working full time. Adjust the balance using the skills matrix above.

WeekFocusConcrete output
1Coding patterns: arrays, hash maps, two pointers, sliding window, intervals15-20 problems, timed, out loud
2Coding: trees, graphs, BFS/DFS, heaps, topological sort15-20 problems; one concurrency problem
3Databases: indexes, EXPLAIN, isolation levels, locking, shardingDesign schemas and indexes for 3 apps; explain 5 query plans
4API design: REST vs gRPC, pagination, idempotency, versioning, errorsWrite full API specs for 3 prompts (orders, tasks, chat)
5System design: caching, queues, outbox, rate limiting, failure modes4 timed designs, each with a "hardest part" deep dive
6Mock interviews and behavioral2-3 mocks; 6-8 STAR stories including one incident and one migration

For behavioral prep, backend interviewers lean on incident stories: an outage you debugged, a migration you led, a time you pushed back on a design. The behavioral interview guide shows how to structure them.

Two habits make the plan work. Write real SQL and real API specs, because "I would add an index" is weaker than the exact index and the query it serves. And end every practice design by asking what breaks first at 10x traffic.

From API pagination edge cases to isolation-level follow-ups, a backend loop tests a lot of recall under time pressure. TechScreen runs invisibly during your screen share and gives real-time help across coding, API design, and system design rounds. Claim 3 free tokens and try it on a mock backend interview first.

Get started free →

Frequently Asked Questions

What is asked in a backend engineer interview?

Most backend loops include one or two coding rounds on data structures and algorithms, a system design round, and a behavioral round. Many companies add a backend-specific round: designing a REST or gRPC API, modeling a database schema, or debugging a service. Expect questions on indexing, transactions and isolation levels, caching strategy, message queues, idempotency, and how a service behaves when a dependency fails or slows down.

How is a backend interview different from a general software engineer interview?

The coding rounds are often identical, but backend candidates are graded harder on data and failure. Interviewers expect you to choose a database and justify it, design indexes for real queries, explain what happens under concurrent writes, and handle retries without double-processing. System design carries more weight, and many companies add an API design or schema design round that a generalist or frontend loop would not include.

Should I use REST or gRPC in an API design interview?

Default to REST with JSON for public, browser-facing, or partner APIs, because every client can call it and it is easy to debug. Choose gRPC for internal service-to-service calls where you control both sides and want strict Protocol Buffers contracts, generated clients, lower payload size, and streaming over HTTP/2. Say why you chose one in a sentence. Interviewers care about the reasoning far more than the pick itself.

Which database topics come up most in backend interviews?

The recurring topics are B-tree indexes and composite index column order, covering indexes, ACID transactions, isolation levels and the anomalies each one allows, optimistic versus pessimistic locking, replication lag and read-your-writes consistency, sharding keys and hot partitions, and when to choose a relational database over a key-value or document store. Being able to read a query plan and explain why a query is slow is a strong signal.

How long does it take to prepare for a backend engineer interview?

Working engineers usually need four to eight weeks at eight to twelve hours per week. A practical split is roughly 40 percent coding practice, 35 percent system design and API design, 15 percent databases and concurrency fundamentals, and 10 percent behavioral stories. Senior and staff candidates should shift more time to system design and to stories about incidents, migrations, and cross-team technical decisions.

Do backend engineers still need to do LeetCode in 2026?

At most large tech companies, yes. Google, Meta, Amazon, and Microsoft still run algorithm rounds for backend roles, usually at medium difficulty. Some companies, including Stripe and many startups, replace puzzles with practical rounds such as building against an API or debugging an existing codebase. Plan for both: cover the core patterns well rather than grinding hundreds of problems, then practice building small working services.

What is idempotency and why do interviewers ask about it?

An operation is idempotent if performing it several times has the same effect as performing it once. Interviewers ask because networks fail and clients retry, so a payment or order endpoint without idempotency can charge a customer twice. The standard answer is a client-generated idempotency key stored with the result, checked before processing, and enforced with a unique constraint so concurrent duplicates cannot both succeed.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →