← All articles
12 min read

Debugging Interview Guide 2026: How to Ace Live Debugging

A debugging interview tests how you find a bug in code you have never seen. This guide gives you a repeatable 6-step protocol and walks it through a realistic multi-file bug, line by line.

A debugging interview is a live round where you get an unfamiliar codebase with a known bug and have to reproduce it, find the root cause, and ship a correct fix while explaining your reasoning. Interviewers grade your method more than your speed. The candidates who pass follow a repeatable process: reproduce, map, hypothesize, probe, fix, and guard against regression.

This guide gives you that process as a 6-step protocol, then runs it on a realistic bug spread across four files.

Key Takeaways

  • Debugging rounds score process. A clear trail of stated hypotheses and evidence beats a lucky fix you cannot explain.
  • Get a failing signal first. A failing test or a reliable reproduction in the first 10 minutes sets up everything after it.
  • Read structurally before reading deeply. Trace one path from the entry point to the symptom and ignore the rest of the codebase.
  • Every print, log, or breakpoint should answer one specific question you said out loud first.
  • A symptom suppressed in the wrong file is the most common failing answer. Fix the cause, then add a regression test.
  • Codebase-based and AI-enabled rounds are growing because they are harder to shortcut than known algorithm puzzles.

Why Are Debugging Rounds Growing in 2026?

Debugging rounds are growing because they measure the job more directly than algorithm puzzles do. Most engineering time goes into reading, changing, and fixing existing code, and a bug in a real codebase tests that work in a way a fresh LeetCode problem cannot.

AI is the second driver. Classic algorithm problems are widely documented online, and AI models solve many of them quickly. That has pushed companies toward formats built around a specific codebase, where the hard part is understanding context. Meta has added an AI-enabled coding round to its loop, as our Meta interview breakdown describes, and our AI-enabled coding interview guide covers how those rounds work. In those formats the skills overlap heavily with debugging: reading generated or existing code critically, testing it, and catching what is wrong.

Stripe's Bug Squash round is a well-known pure debugging format. Our Stripe interview process breakdown describes it in detail: an unfamiliar codebase, a known bug, and roughly an hour to fix it. Similar rounds appear under names like "code reading," "fix the failing test," or "production incident" at some companies. For a wider view of the shift, see are coding interviews getting harder.

What Do Interviewers Score in a Debugging Round?

Interviewers score four things: orientation speed, hypothesis quality, tool fluency, and fix quality. Finding the bug matters, but how you got there carries most of the signal.

SignalStrong candidateWeak candidate
OrientationFinds the entry point and the path to the symptom in minutesReads files top to bottom, alphabetically
HypothesesStates 2-3 ranked guesses, each tied to evidenceChanges code to "see what happens"
Tool useRuns one test, sets targeted breakpoints, searches the repoAdds 20 prints, reads none carefully
Fix qualitySmallest change at the root cause, plus a regression testSpecial-case patch where the symptom appears
CommunicationNarrates what is known, unknown, and nextGoes silent for long stretches

The common thread is discipline. Our guide to what interviewers look for in coding interviews covers the general version of this rubric. Debugging rounds simply make it visible.

The 6-Step Debugging Protocol

The protocol is a fixed sequence you can run on any bug in any language: Reproduce, Map, Hypothesize, Probe, Fix, Guard. Each step has a goal, a time budget for a 45 to 60 minute round, and a line you can say out loud so the interviewer can follow you.

StepGoalBudgetWhat to say
1. ReproduceGet a reliable failing signal5-10 min"First I want to see it fail on demand."
2. MapTrace the path from entry point to symptom5-10 min"Here is how a request flows to where the bug shows up."
3. HypothesizeList 2-3 ranked causes2-3 min"My top guess is X because of Y. Second is Z."
4. ProbeRun the cheapest check that splits the hypotheses10-15 min"If X is right, this value will be wrong here."
5. FixSmallest change at the root cause5-10 min"The cause is here, so the fix belongs here."
6. GuardRerun, add a regression test, check for siblings5 min"Let me confirm nothing else relies on the old behavior."

Step 1: Reproduce

Reproduction means turning a vague bug report into a signal you can trigger in seconds. Run the existing test suite first. If a test already fails, you have your signal. If not, write the smallest test or script that shows the wrong behavior.

Ask the interviewer for the expected behavior if the report is ambiguous. "Totals are wrong" is not testable. "The total should equal subtotal plus tax" is.

Step 2: Map

Mapping means finding the one path through the code that matters. Start at the entry point (a request handler, a CLI command, a public method) and follow calls toward the code that produces the wrong output. Skip everything else.

Step 3: Hypothesize

A hypothesis is a specific, testable claim about the cause, such as "the discount is applied twice." Write two or three down, ranked by likelihood and by how cheap they are to check. Vague hypotheses like "something with state" cannot be tested, so sharpen them first.

Step 4: Probe

A probe is the cheapest experiment that confirms or kills a hypothesis. Good probes split the search space in half. Pick a spot between the input and the symptom, check the value there, and you know which half holds the bug.

Step 5: Fix

Fix the cause, not the place where the symptom shows. If a value is corrupted in one module and displayed wrong in another, the fix belongs where it gets corrupted. Keep the change small enough to explain in one sentence.

Step 6: Guard

Guarding means proving the fix and protecting it. Rerun the failing test, run the full suite, and add a test that would have caught this bug. Then search for the same pattern elsewhere, because bugs often have siblings.

How Do You Read Unfamiliar Code Fast?

You read unfamiliar code fast by reading less of it. In an interview you need the shape of the system and one execution path, not a full understanding of every file.

A practical order of operations:

  1. Read the README, the directory tree, and the test file names. Tests are often the best documentation, because they show how each module is meant to be called.
  2. Find the entry point for the failing behavior. Search for the function name, route, or error message from the bug report.
  3. Read signatures before bodies. Function names, parameters, and return types tell you the data flow.
  4. Follow the data, not the call stack. Track the one value that ends up wrong and every place that reads or writes it.
  5. Note surprises in a scratch file: mutation of inputs, global state, caching, implicit type conversions, and time or randomness.

Text search is your fastest tool. Searching for every assignment to a field, for example rg "unit_price\s*=", often finds a state bug faster than reading. If the code is in a language you rarely write, say so and lean on the tests. Reading code is a far more portable skill than writing it, and interviewers know that.

Tests, Logs, and Breakpoints: Which Tool When?

Each debugging tool answers a different kind of question. Picking the right one is part of what interviewers grade.

ToolBest forInterview tip
Unit testReproducing and verifying the bugRun one test at a time, for example pytest -x -k total
Print or logSeeing a value change across several callsLabel every print with location and variable name
BreakpointInspecting full state at one momentIn Python, breakpoint() drops you into pdb
Stack traceFinding where an exception startedRead from the bottom frame of your own code, not the library
git bisectFinding which commit introduced a regressionMention it if history exists; see the git bisect docs
Repo searchFinding every reader and writer of a valueSearch assignments, not just usages

A useful rule: use prints when you need to see a sequence over time, and use a breakpoint when you need many variables at one point. Use tests to bracket the whole session, failing at the start and passing at the end.

Ask early whether you can use your own editor and debugger. Many debugging rounds allow it, and our article on looking things up during a coding interview covers what is usually acceptable in terms of docs and references.

Debugging rounds go wrong when you lose the thread under pressure. TechScreen runs invisibly on Zoom, Google Meet, and Teams screen shares and gives you real-time hints on where to look next in unfamiliar code. Start with 3 free tokens.

Get started free →

Worked Example: A Multi-File Pricing Bug

This example runs the full protocol on a small Python checkout service with four files. The bug report reads: "For bulk orders, the order total is lower than the subtotal shown plus tax."

The relevant code:

# tiers.py
TIERS = [(100, 0.10), (50, 0.05), (10, 0.02)]

def discount_for(qty):
    for threshold, rate in TIERS:
        if qty >= threshold:
            return rate
    return 0.0
# pricing.py
from tiers import discount_for

def line_total(item):
    rate = discount_for(item["qty"])
    if rate:
        item["unit_price"] = round(item["unit_price"] * (1 - rate))
    return item["unit_price"] * item["qty"]
# tax.py
TAX_RATE = 0.08

def compute_tax(cart):
    return round(cart.subtotal() * TAX_RATE)
# cart.py
from pricing import line_total
from tax import compute_tax

class Cart:
    def __init__(self):
        self.items = []

    def add(self, sku, qty, unit_price):
        self.items.append({"sku": sku, "qty": qty, "unit_price": unit_price})

    def subtotal(self):
        return sum(line_total(item) for item in self.items)

    def summary(self):
        subtotal = self.subtotal()
        tax = compute_tax(self)
        return {"subtotal": subtotal, "tax": tax, "total": self.subtotal() + tax}

Reproduce

No existing test covers this, so write one. Prices are in cents.

# test_cart.py
from cart import Cart

def test_total_equals_subtotal_plus_tax():
    cart = Cart()
    cart.add("WIDGET", qty=120, unit_price=1000)
    s = cart.summary()
    assert s["total"] == s["subtotal"] + s["tax"]

The test fails. The summary reports a subtotal of 108,000, tax of 7,776, and a total of 95,256. Two things look wrong: the total is far too low, and 8% of 108,000 should be 8,640, not 7,776. Say that out loud. Two wrong numbers from one input is a strong clue.

Map

The path is Cart.summary to Cart.subtotal to line_total to discount_for. Tax goes through compute_tax, which calls cart.subtotal() again. Note that subtotal() runs three times during one summary() call.

Hypothesize

Three candidates, ranked:

  1. line_total changes state, so repeated calls return different values. This explains both wrong numbers.
  2. The tier boundaries in tiers.py are wrong, giving the wrong rate for 120 units.
  3. Tax rounding is off in tax.py.

Hypothesis 3 cannot explain the low total, and hypothesis 2 would give a consistent wrong answer rather than different answers within one call. Hypothesis 1 goes first.

Probe

The cheapest probe that splits these is to call subtotal() twice and compare:

cart = Cart()
cart.add("WIDGET", qty=120, unit_price=1000)
print("first:", cart.subtotal(), "unit:", cart.items[0]["unit_price"])
print("second:", cart.subtotal(), "unit:", cart.items[0]["unit_price"])

The output shows first: 108000 unit: 900 then second: 97200 unit: 810. The unit price shrinks on every call, which confirms hypothesis 1. A quick check that discount_for(120) returns 0.10 rules out hypothesis 2. Searching for unit_price\s*= across the repo shows the only write outside Cart.add is the line in pricing.py.

Fix

The symptom shows up in cart.py, but the cause is in pricing.py: line_total writes the discounted price back into the shared item dict. The fix is to compute the price locally.

# pricing.py
from tiers import discount_for

def line_total(item):
    rate = discount_for(item["qty"])
    unit_price = round(item["unit_price"] * (1 - rate))
    return unit_price * item["qty"]

A weaker fix would cache the subtotal inside summary() so it is computed once. That makes the test pass but leaves line_total corrupting data for every other caller. Interviewers catch this and score it as a symptom patch.

Guard

Rerun the test: subtotal 108,000, tax 8,640, total 116,640. Then add a regression test that targets the actual cause:

def test_line_total_does_not_mutate_item():
    item = {"sku": "WIDGET", "qty": 120, "unit_price": 1000}
    line_total(item)
    line_total(item)
    assert item["unit_price"] == 1000

Finish by naming the class of bug: a function with a hidden side effect on shared input. Mention that you would look for other pricing helpers that write back into item dicts. That one sentence shows you think about the codebase, not just the ticket.

How Should You Communicate Hypotheses?

Communicate in short loops: what you know, what you suspect, what you will check, and what the result means. Narration lets the interviewer give credit for reasoning even before you find the bug, and it invites hints.

Phrases that work well in a live debugging interview:

  • "The symptom is X. The expected behavior is Y. Let me confirm that with you."
  • "I have three guesses. The first explains both wrong numbers, so I will check it first."
  • "If this hypothesis is right, this value should be 900 here. Let me look."
  • "That ruled out the tier logic. Moving on to state changes between calls."
  • "I could patch this in cart.py, but the cause is upstream, so I will fix it there."

Keep silences under about a minute. If you need to read quietly, say so: "Give me 30 seconds to read this function." Our guide on how to think out loud in a coding interview has more scripts, and the pair programming interview guide covers how to work with an interviewer who acts as a teammate.

If you hit a wall, go back to the protocol rather than guessing. Restate what you have ruled out and pick a new split point. The what to do when stuck guide covers recovery moves in more detail.

What Are the Most Common Debugging Interview Mistakes?

Most failed debugging rounds come from the same handful of habits. Each one has a simple fix.

  • Fixing before reproducing. Without a failing signal you cannot prove the fix worked. Get the test red first.
  • Reading the whole codebase. Thirty minutes of reading leaves no time to probe. Map one path.
  • Shotgun prints. Twenty unlabeled prints create noise. Add one probe per hypothesis and label it.
  • Patching the symptom. A special case where the wrong number appears is the classic failing answer.
  • Silence. Long quiet stretches hide good reasoning. Narrate in short loops.
  • Ignoring the tests. Existing tests show intended behavior and often point straight at the bug.
  • Changing several things at once. If two changes fix it, you do not know which one did. Change one thing, rerun.
  • Skipping the regression test. A fix with no test looks unfinished, even when it is correct.

To practice, take a small open-source project in your main language, plant a bug in a function two or three calls away from where it shows, and run the six steps against a 45-minute timer. Do it again in a language you read but rarely write. Five or six reps make the protocol automatic, so the pressure goes into thinking instead of remembering what to do next.

Want backup on interview day? TechScreen is an invisible AI interview assistant that stays hidden during screen shares on Zoom, Google Meet, Teams, HackerRank, and CoderPad, so you can check a hypothesis or get unstuck in a live debugging round without anyone seeing it. Try it with 3 free tokens.

Get started free →

Frequently Asked Questions

What is a debugging interview?

A debugging interview is a technical round where you get an existing codebase with one or more known bugs and have to reproduce, diagnose, and fix them while the interviewer watches. It replaces or supplements the classic algorithm round. Interviewers score your process: how fast you orient yourself in unfamiliar code, whether you form testable hypotheses, how you use tests, logs, and debuggers, and whether your fix addresses the root cause instead of hiding the symptom.

How do I prepare for a live debugging interview?

Practice on real code, not puzzles. Clone a small open-source project in your main language, have a friend or a script plant a bug, then fix it against a 45-minute timer while narrating your reasoning out loud. Repeat with a second language you read but rarely write. Get fluent with your language's test runner flags, a debugger or breakpoint(), and text search across a repository, because those tools carry most of the work in a real round.

Can I use a debugger or print statements in a debugging interview?

Yes, in most debugging rounds. Interviewers generally expect you to run the code and gather evidence, and both print statements and breakpoints are normal. What they watch is intent: each probe should test a specific hypothesis you stated first. Scattering prints everywhere without a question in mind reads as guessing. If the environment limits tools, ask at the start what you can run, then adapt.

What if I cannot find the bug before time runs out?

Partial credit is real in debugging rounds, because the process is what gets scored. If time is short, state the narrowest region where you believe the bug lives, the evidence that put it there, the hypotheses you ruled out, and the next probe you would run. A candidate who has cut the search space to one function with clear reasoning often scores better than one who lands a lucky fix without explaining it.

Are debugging interviews harder than LeetCode interviews?

They are different rather than strictly harder. Debugging rounds reward working-engineer skills: reading code quickly, using tools, and reasoning from evidence. Candidates who have shipped production code often find them more natural than algorithm puzzles. Candidates who prepared mostly with LeetCode tend to struggle, because memorized patterns do not help when the problem is a state bug spread across three files they have never opened.

Which companies use debugging rounds in 2026?

Stripe's Bug Squash round is a well-known example, and other companies run similar codebase-based rounds. The format appears to be spreading as companies look for interviews that hold up when candidates have AI tools, since finding a real bug in a real codebase is harder to shortcut than a known puzzle. Ask your recruiter directly whether a debugging or code-reading round is part of your loop.

Ready to use AI assistance in your next interview?

TechScreen is the invisible AI assistant trusted by engineers interviewing at Google, Meta, Amazon, and hundreds of other companies. Start with 3 free tokens — no credit card required.

Ace your next interview →