← All articles
13 min read

Payment System Design Interview: Idempotency and Ledgers 2026

A payment system design answer lives or dies on correctness. This guide covers idempotency keys, double-entry ledgers, and reconciliation, then walks through exactly what happens when a charge times out.

Payment system design is the interview question where correctness matters more than scale: your system must never charge a customer twice, never lose a payment it already took, and always be able to explain every cent. A strong answer rests on three ideas: idempotency keys so retries are safe, a double-entry ledger so money is always accounted for, and reconciliation so you catch the errors the first two miss. If you can walk an interviewer through what happens when a charge times out, you are ahead of most candidates.

Key Takeaways

  • A timeout from the payment provider means "unknown," not "failed." Marking it failed and letting the user retry is the classic double-charge bug.
  • Use two layers of idempotency: a client key on your API, and a stable key your service sends to the provider.
  • Exactly-once delivery is not possible across a network. Exactly-once effects are, through at-least-once delivery plus deduplication.
  • A double-entry ledger is append-only, every transaction sums to zero, and balances are derived from entries.
  • Reconciliation against the provider's settlement reports is not optional. It is the safety net for every bug you did not anticipate.
  • Keep raw card data out of your systems entirely by tokenizing through the provider, which shrinks your PCI DSS scope.

What Does a Payment System Design Interview Actually Test?

It tests whether you design for failure first. Most system design questions let you trade some correctness for availability. Payments do not, so interviewers watch for you asking "what if this call fails halfway?" at every arrow.

The question appears as "design a payment system," "design Stripe," "design checkout," or "design a wallet." It is a staple at fintech companies, and our Stripe interview guide lists money movement and idempotency among likely system design topics. Companies like Coinbase and Robinhood ask close cousins of it.

The general framework from our guide on how to ace a system design interview still applies: clarify requirements, sketch the high-level design, go deep on the hard parts, then cover scaling and failure. For payments, the "hard parts" are almost entirely about correctness.

Requirements and Actors

Start by naming the actors and agreeing on scope. A payment system is a set of services that moves money between parties through external providers. You almost never move money yourself.

ActorRoleWhat you control
CustomerPays for goods or servicesNothing, but you hold their payment token
Merchant (or seller)Receives the moneyTheir balance and payout schedule
Your payment serviceOrchestrates the payment, owns stateEverything inside your boundary
PSP (payment service provider)Stripe, Adyen, Braintree; talks to card networksOnly the requests you send
Card networks and issuing banksApprove or decline, move fundsNothing
LedgerYour record of who owes whatFully owned by you

Clarify which side you are on. Designing a marketplace's payment system means you call a PSP. "Design Stripe" means you are the PSP, which adds card network integration, merchant onboarding, and fraud at a much larger scale. Ask in the first two minutes.

Functional requirements

  • Accept a payment for an order (authorize and capture, or a single charge).
  • Support refunds, full and partial.
  • Record every money movement in a ledger.
  • Pay out merchants on a schedule.
  • Notify merchants and internal systems of payment status changes.

Non-functional requirements

  • No double charges and no lost payments. This beats every other requirement.
  • Auditability. Every balance must be explainable from history.
  • Availability is important, but a payment can be delayed safely. It cannot be duplicated safely.
  • Scale is usually modest. Even large merchants process far fewer payments per second than a social feed serves reads. Do a quick estimate, then spend your time on correctness.

Payment Flow End to End

A payment flow is the sequence of state changes from "customer clicks pay" to "merchant has the money." Here is the version you should be able to draw from memory:

  1. The client collects card details in the PSP's hosted fields or SDK. Raw card numbers go straight to the PSP, which returns a token. Your servers never see the card number.
  2. The client calls your API: POST /payments with the order ID, amount, currency, token, and an Idempotency-Key header.
  3. The payment service checks the idempotency store. New key: it inserts a payment record in state CREATED.
  4. The payment service calls the PSP with its own idempotency key (derived from the payment ID) and moves the record to PENDING.
  5. The PSP returns success or decline. The service moves the payment to SUCCEEDED or FAILED.
  6. On success, the service writes balanced ledger entries in the same database transaction as the state change, or via an outbox event.
  7. The service publishes a payment.succeeded event. Downstream consumers (order service, email, merchant webhooks) react to it.
  8. Separately, the PSP sends webhooks. Your webhook handler deduplicates by event ID and reconciles state with what you already know.
  9. Nightly, a reconciliation job compares your ledger to the PSP's settlement report.

Model the payment as an explicit state machine. Allowed transitions should be enforced in code and in the database with a conditional update such as UPDATE payments SET status = 'SUCCEEDED' WHERE id = ? AND status = 'PENDING'. If zero rows change, another worker already moved it, and you stop.

CREATED -> PENDING -> SUCCEEDED -> (PARTIALLY_)REFUNDED
                   \-> FAILED
                   \-> UNKNOWN -> SUCCEEDED | FAILED   (resolved by retry or status query)

The UNKNOWN state is the part most candidates leave out, and it is where the interesting discussion happens.

Idempotency and Exactly-Once Effects

An idempotency key is a unique client-generated value that lets the server recognize a retried request and return the original result instead of repeating the operation. It is the single most important mechanism in payment system design.

Per Stripe's API documentation on idempotent requests, Stripe stores the status code and body of the first request made with a key, returns that same result for later requests with the same key (including error responses), compares parameters and errors if a reused key carries different ones, and may prune keys once they are at least 24 hours old. Those are good defaults to copy in your own design.

How the idempotency layer works

def create_payment(request, idem_key):
    record = idem_store.get(idem_key)
    if record:
        if record.request_hash != hash(request.body):
            return error(422, "Idempotency key reused with different parameters")
        if record.status == "IN_PROGRESS":
            return error(409, "Request with this key is still processing")
        return record.saved_response

    inserted = idem_store.insert_if_absent(
        key=idem_key,
        request_hash=hash(request.body),
        status="IN_PROGRESS",
    )
    if not inserted:
        return error(409, "Concurrent request with this key")

    response = process_payment(request)
    idem_store.complete(idem_key, response)
    return response

Three details separate a senior answer from a mid-level one:

  • The insert must be atomic. Use a unique constraint on the key column so two concurrent requests cannot both pass the check. A read-then-write without a constraint has a race.
  • Hash the parameters. A reused key with a different amount is a client bug. Reject it loudly instead of silently returning a stale response.
  • Use a second key downstream. Your service should send the PSP a key derived from your payment ID, such as pay_123:charge. When your worker retries the PSP call after a crash, the PSP deduplicates it.

Exactly-once delivery vs exactly-once effects

Exactly-once delivery across a network is not achievable: a sender cannot tell a lost request from a lost response. What you build instead is exactly-once effects. Deliver messages at least once, and make every consumer idempotent so a duplicate changes nothing.

In practice that means a unique constraint everywhere money moves: on idempotency keys, on PSP webhook event IDs, and on ledger entries per (payment ID, entry type). The database becomes your deduplication engine. If you want a comparison point, the notification system design guide uses the same at-least-once plus dedupe pattern for messages, where the stakes are lower.

Payment follow-ups come fast: "what if the worker crashes after the PSP call?" TechScreen gives you real-time structure for answers like this during live system design rounds, invisible on Zoom, Meet, and Teams screen shares. Start with 3 free tokens, no credit card.

Get started free →

Worked Example: The Timed-Out Charge

This walkthrough shows exactly how the design above behaves when a PSP call times out. Interviewers love it because it exercises every layer at once.

Setup. A customer pays $80 for order ord_42. The client sends POST /payments with Idempotency-Key: k-7f3a.

Step 1: Happy start. The service inserts key k-7f3a as IN_PROGRESS, creates payment pay_123 in CREATED, moves it to PENDING, and calls the PSP with key pay_123:charge.

Step 2: The timeout. The PSP authorizes and captures the card, but the response never arrives. After the configured timeout, the HTTP client throws.

Step 3: The wrong move. A naive service marks pay_123 as FAILED and returns an error. The customer clicks "Pay" again, a new payment pay_124 is created with a new PSP key, and the card is charged a second time. This is the bug the whole question is designed to surface.

Step 4: The right move. The service moves pay_123 to UNKNOWN, writes no ledger entries, and tells the client the payment is processing. The client shows a spinner and polls GET /payments/pay_123 instead of resubmitting.

Step 5: The client retries anyway. A mobile app on a flaky network resends the original request with the same key k-7f3a. The idempotency store returns 409 (still processing) or the same pay_123 reference. No new payment is created.

Step 6: Resolution. A recovery worker scans for UNKNOWN payments older than a few seconds. It either retries the PSP call with the same key pay_123:charge, which returns the original success without charging again, or it queries the PSP for the payment's status. Either way, it learns the charge succeeded.

Step 7: State and ledger. The worker runs a conditional update from UNKNOWN to SUCCEEDED and, in the same transaction, inserts the ledger entries. A unique constraint on (pay_123, capture) guarantees the entries exist once.

Step 8: The webhook races in. The PSP's charge.succeeded webhook arrives. The handler records the event ID (deduplicated), sees pay_123 is already SUCCEEDED, and does nothing. Had it arrived first, it would have done the resolution itself, and the worker's conditional update would have changed zero rows.

Step 9: The safety net. That night, reconciliation matches pay_123 against the PSP's settlement report. If some bug had left it FAILED while the PSP shows a capture, reconciliation flags it for automatic repair or human review.

Failure pointNaive outcomeCorrect handling
PSP response times outMarked failed, user charged twice on retryMark UNKNOWN, resolve by retry with same PSP key or status query
Client resends requestNew payment createdSame idempotency key returns existing payment
Worker crashes after PSP callPayment stuck in PENDING foreverRecovery job sweeps stale PENDING and UNKNOWN rows
Webhook and worker both resolveLedger written twiceConditional state update plus unique ledger constraint
Silent bug in any of the aboveBooks drift, nobody noticesNightly reconciliation flags the mismatch

Double-Entry Ledger Design

A double-entry ledger is an append-only record in which every transaction is split into debits and credits that sum to zero. It makes errors visible: if money appears from nowhere, a transaction will not balance.

Schema

CREATE TABLE accounts (
  id            BIGINT PRIMARY KEY,
  owner_type    TEXT NOT NULL,
  owner_id      TEXT NOT NULL,
  account_type  TEXT NOT NULL,
  currency      CHAR(3) NOT NULL
);

CREATE TABLE ledger_transactions (
  id            BIGINT PRIMARY KEY,
  payment_id    TEXT NOT NULL,
  kind          TEXT NOT NULL,
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
  UNIQUE (payment_id, kind)
);

CREATE TABLE ledger_entries (
  id              BIGINT PRIMARY KEY,
  transaction_id  BIGINT NOT NULL REFERENCES ledger_transactions(id),
  account_id      BIGINT NOT NULL REFERENCES accounts(id),
  amount_minor    BIGINT NOT NULL,
  direction       TEXT NOT NULL CHECK (direction IN ('debit', 'credit'))
);

Example postings

Using the $80 payment, with a hypothetical $2.60 PSP fee and a 10% platform commission on a marketplace:

AccountDebitCredit
PSP receivable (money the PSP owes us)77.40
PSP fees expense2.60
Merchant payable (owed to the seller)72.00
Platform revenue8.00
Total80.0080.00

Rules worth stating out loud

  • Store money as integers in minor units (cents), with a currency code. Never floats.
  • Never update or delete entries. A refund is a new transaction that reverses the original postings. This gives you a full audit trail.
  • Derive balances from entries. Cache balances in a separate table for speed if needed, but treat entries as the source of truth and rebuild the cache from them.
  • Enforce the zero-sum rule at write time, inside the transaction, and verify it with a periodic job.
  • One ledger transaction per business event, keyed by (payment ID, kind), so retries cannot post twice.

For storage, a relational database with strong transactional guarantees is the safe default. If the interviewer asks about scale, shard by account or merchant and keep each ledger transaction within one shard where possible. Our SQL interview questions guide covers the constraint and transaction fundamentals interviewers will probe here.

Reconciliation and Retries

Reconciliation is the process of comparing your internal records against an external source of truth, usually the PSP's settlement or balance reports, and resolving every difference. It exists because no amount of careful design eliminates every bug, outage, or manual dashboard action.

A typical pipeline:

  1. Ingest the PSP's settlement report daily (most providers expose these through an API or file export).
  2. Match each PSP record to an internal payment by your reference ID, then compare amount, currency, status, and fee.
  3. Classify mismatches: missing internally, missing at the PSP, amount differs, status differs.
  4. Auto-resolve known patterns, such as a payment stuck in UNKNOWN that the PSP shows as captured.
  5. Queue the rest for a human finance or operations review, with alerting when the unmatched count or value crosses a threshold.

Retry policy

Retries are where idempotency pays off, but they still need limits:

  • Retry only on errors that are safe and plausibly transient: timeouts, connection errors, and server-side 5xx. Do not retry a card decline.
  • Use exponential backoff with jitter so a PSP outage does not turn into a thundering herd when it recovers.
  • Cap attempts, then move the payment to a dead-letter queue for manual handling.
  • Always reuse the same PSP idempotency key across retries of the same operation. A new key per attempt defeats the purpose.

Protecting your own API from abusive retry loops is a separate problem; the rate limiter design guide covers per-client limits you can put in front of the payment endpoint.

Security and Compliance Basics

You do not need an auditor's knowledge, but you should cover a few points without being asked.

  • PCI DSS. The Payment Card Industry Data Security Standard governs anyone who stores, processes, or transmits card data. The current version is v4.0.1, and according to the PCI Security Standards Council, the requirements that were future-dated in v4.0 became mandatory on March 31, 2025.
  • Tokenization. Collect cards through the PSP's hosted fields or SDK so raw card numbers never reach your servers. This shrinks your compliance scope dramatically, which is the main reason merchants use PSPs at all.
  • Never store security codes. PCI DSS prohibits storing sensitive authentication data such as the CVV after authorization.
  • Strong customer authentication. In the EU and UK, card payments often require two-factor checks such as 3-D Secure. Your flow needs a state for "waiting on customer authentication."
  • Webhook authenticity. Verify the PSP's webhook signature before trusting any event, and reject stale timestamps to block replays.
  • Access and audit. Encrypt data in transit and at rest, keep secrets in a managed vault, and log every manual balance adjustment with who made it and why.

Common Follow-Up Questions

Expect at least two of these after your main design. Short, confident answers win.

  • "How do you handle refunds?" A refund is its own payment-like object with its own idempotency key and state machine, and it posts reversing ledger entries. Partial refunds must never exceed the captured amount, so check against ledger-derived totals inside a transaction.
  • "What if your database is down?" Fail closed. Reject new payments rather than calling the PSP without recording state, because a charge you cannot record is worse than a delayed one.
  • "How do you pay out merchants?" A scheduled job sums each merchant's payable balance from the ledger, creates a payout with its own idempotency key, and posts entries moving funds from merchant payable to cash.
  • "Why not use a distributed transaction across your DB and the PSP?" You cannot: the PSP is an external HTTP API. That is precisely why you need a state machine, idempotency keys, and reconciliation.

If you are aiming for a senior or staff loop, add a word on operations: dashboards for stuck UNKNOWN payments, alerts on reconciliation drift, and runbooks for PSP outages. The staff engineer interview guide explains why that operational lens matters at higher levels.

How to Practice This Question

Practice the timed-out charge walkthrough out loud, and draw the state machine and ledger postings from memory. Then take each arrow and ask, "what if this fails before, during, and after?" For a routine, see our guide on how to practice system design interviews, and if you are interviewing for backend roles broadly, the backend engineer interview guide shows where payments questions fit in the full loop.

Want a second brain in the room when the interviewer asks "and what happens if the webhook arrives first?" TechScreen is a real-time AI interview assistant that stays invisible during screen shares. Try it with 3 free tokens, no credit card required.

Get started free →

Frequently Asked Questions

What is the most important concept in a payment system design interview?

Correctness under failure. Interviewers care less about raw throughput and more about whether a customer can ever be charged twice, or whether money can appear or vanish from your books. The three tools that answer that are idempotency keys on every write that moves money, a double-entry ledger where every transaction balances to zero, and reconciliation against the payment provider's records to catch anything the first two missed.

How do idempotency keys prevent duplicate charges?

The client generates a unique key, usually a UUID, and sends it with the payment request. The server stores the key with the request parameters and the final response. If the same key arrives again, the server returns the stored response instead of charging again. Stripe's API works this way and rejects a reused key whose parameters differ from the original. Your service should also pass its own stable key to the payment provider so retries downstream are safe too.

Why do payment systems use double-entry ledgers?

A double-entry ledger records every movement of money as at least two entries, a debit and a credit, that sum to zero. This makes errors detectable: if any transaction does not balance, or account totals drift from expected values, something is wrong and you can find it. Entries are append-only, so corrections are new reversing entries rather than edits, which gives auditors and engineers a complete history of how every balance was reached.

What should happen when a call to the payment provider times out?

Treat the outcome as unknown, not failed. The provider may have charged the card even though you never saw the response. Mark the payment as pending or unknown, never as failed, and then resolve it by retrying with the same provider idempotency key or by querying the provider for the payment's status. Only write the ledger entries once you know the real outcome, and let reconciliation catch anything that slips through.

Can you get exactly-once payment processing in a distributed system?

Not exactly-once delivery, since networks drop and duplicate messages. What you can build is exactly-once effects: messages may be delivered more than once, but each one changes state only once. You get there with at-least-once delivery plus deduplication, using idempotency keys at the API, unique constraints in the database, and idempotent consumers for webhooks and queue events. Saying this distinction clearly is a strong signal in interviews.

How much PCI DSS knowledge do I need for a system design interview?

You need the basics, not an auditor's depth. Know that raw card numbers should never touch most of your services, that tokenization through the payment provider's hosted fields shrinks your compliance scope, that card security codes must never be stored after authorization, and that PCI DSS v4.0.1 is the current standard. Mentioning that you would isolate any card data environment and keep it minimal is usually enough.

Is designing Stripe the same as designing a payment system?

Not quite. Designing a payment system usually means a merchant or marketplace that calls a provider like Stripe or Adyen. Designing Stripe means you are the provider, so you also own card network connections, merchant onboarding, payouts, and fraud scoring at much larger scale. The core correctness ideas are identical, so clarify which side of the integration you are building in the first two minutes.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →