The difference between one good fix and three confident guesses.
debug-this.txt
Something is broken. Here is everything I have.
What I expected: <expected>
What actually happens: <actual>
Exact error (full text, not a summary):
<paste>
The file involved:
<paste the whole file>
Before suggesting a fix, tell me in one sentence what you think the
root cause is. If you are guessing, say so and tell me what to check
or log to confirm it.
Never open with "it doesn't work". Models will happily invent a plausible cause for a symptom you haven't described, and you'll spend the next twenty minutes disproving it.
Error, full text:
<paste>
The code that produced it:
<paste>
In this order:
1. What this error means in general — in plain English, no jargon.
2. What specifically triggered it in MY code, pointing at the line.
3. The fix.
4. What I could have checked to spot this myself next time.
If you cannot tell from what I've given you, say what else you need.
Step 4 is the one that compounds. Errors repeat; the second time you meet this one you want to fix it in thirty seconds, not thirty minutes.
Something goes wrong somewhere between <the start> and <the result>.
I don't know where.
Don't guess at a fix. Instead, give me a numbered list of checks that
splits the path roughly in half each time — the fastest sequence to
narrow down where the data stops being what I expect.
For each check, tell me exactly what to add or run, and what result
would point which way.
Guessing scales badly. Halving scales beautifully: ten steps narrows a thousand lines to one.
Before I run this, audit your own answer.
List every package, module, API, function and config option you used
that is not something I wrote. For each one:
- Is it real, and is it still the current way to do this?
- Which version are you assuming?
- If you are not confident it exists as you described, say so plainly.
I would rather have "I'm not sure" than a package name that 404s.
Models produce plausible package names the way they produce plausible sentences. Asking for a confidence pass costs one message and catches most of them.
Something goes wrong <sometimes / for some users / occasionally>:
<describe what you see>
Help me make it reproducible. Ask me questions if you need to, then:
- What differs between the times it happens and the times it doesn't?
Suggest specific things to compare — input, timing, data state,
browser, network, order of actions.
- What should I log to capture the conditions when it next occurs?
- Is there a way to force the conditions rather than waiting?
Don't suggest a fix. I want a reliable reproduction first.
Intermittent nearly always means timing, state left over from a previous action, or one specific piece of data. Those three cover most of them.
This fails: <paste the code and the failure>
Help me reduce it to the smallest thing that still fails. Give me an
order to remove things in — largest and least likely to be involved
first — checking after each removal that it still fails.
I want to end up with something under 20 lines that reproduces it.
If removing something makes the failure stop, that's the interesting
part — tell me what to conclude from that.
The reduction usually finds the bug before you finish. Roughly half the time you never need to ask the actual question.
This works: <paste>
This doesn't: <paste>
They should behave the same. Diff them properly — not just the
obvious textual differences, but differences in what they actually do:
different data, different timing, different context, different
assumptions about what exists.
List every difference you find, then rank them by how likely each is
to explain the failure.
Tell me which single one to test first.
Also compare the inputs, not just the code. Identical code with a null in one dataset and not the other is the most common version of this.
I need to find where this goes wrong: <paste the code and describe the
symptom>
Add logging that would identify the failure point. For each line you
add, tell me what a wrong value there would prove.
Log the shape and content of data at each boundary — going in, coming
out — not just "reached here". I need to see where the data stops
being what I expect.
Keep it to <6> log lines. I want the smallest set that narrows it,
not a trace of everything.
Log values, not milestones. "Got here" tells you the path; the actual value tells you what went wrong, and usually why.
Full stack trace: <paste all of it, not just the last line>
The code at the line it points to: <paste, with surrounding context>
Walk me through it:
1. Read from the bottom up — what was the chain of calls?
2. Which is the first frame that's MY code rather than a library?
That's where I should look.
3. What does this error type actually mean?
4. What value was probably not what my code assumed?
5. The fix, and what I could have logged to spot this in seconds.
Pasting only the last line loses most of the information. The chain of calls above it is what tells you how you got into the broken state.
This takes <n> seconds and shouldn't: <paste the code or describe the
operation>
Don't optimise anything yet.
1. How do I measure where the time actually goes? Give me the specific
tool or the timing code to add.
2. Given this code, what are the candidates, ordered by how often each
turns out to be the real cause?
3. Which single measurement distinguishes between them?
I'll come back with numbers. Then we fix the biggest one only.
In server code, check for a query inside a loop before anything else. It's the cause more often than everything else combined.
Works nine times out of ten, and the tenth is the bug.
race-condition.txt
Something intermittent is happening: <describe the symptom and how
often>
Here's the code: <paste>
Look for races specifically:
- Two operations that can overlap and touch the same state.
- A check followed by an action, with a gap between them where things
can change.
- Async operations resolving in a different order than they started.
- A response arriving after the thing that requested it has gone away.
- Two requests both reading, both deciding, both writing.
For each: the exact sequence of events that produces the bug, and the
fix. Then tell me how I could force that sequence to reproduce it
reliably.
"Check then act" is the shape to look for. Anything that reads a value, decides, then writes, has a gap — and under load something will fit into it.
My <app / page / worker> uses more memory the longer it runs, and a
restart fixes it. Stack: <stack>.
Here's the relevant code: <paste>
Look for the usual causes:
- Event listeners added and never removed.
- Timers and intervals never cleared.
- Things pushed into an array or map that's never pruned.
- Closures holding references to large objects.
- Connections, streams or file handles opened without being closed.
- Caches with no eviction.
For each candidate: how it leaks, and how I'd confirm it with a
heap snapshot or a simple counter.
An unbounded cache is the most common one and the least likely to look like a bug. It reads as an optimisation right up to the point the process dies.
A test that fails sometimes is worse than no test.
flaky-test.txt
This test passes sometimes and fails sometimes: <paste the test and
the code under test>
Find the source of the non-determinism:
- Timing — a fixed wait instead of waiting for a condition.
- Shared state between tests, or tests depending on run order.
- Real dates, times or randomness.
- Real network calls.
- Database state not reset between runs.
- Parallel tests touching the same records.
Fix it properly. Do not add a retry or increase a timeout to hide it —
if the code under test has a real race, I want to know that instead.
Retrying a flaky test converts a real bug into an intermittent production incident. The flakiness is often the code, not the test.
Worked in March, broken now, hundreds of commits between.
find-the-commit.txt
This worked at some point and doesn't now. I have <n> commits between
the last known good version and now.
Symptom: <describe>
How I test whether it's broken: <the exact steps or command>
Walk me through bisecting this with git. Include:
- Finding a known-good commit to start from.
- The exact commands, in order.
- What to do when a commit won't build or run — that's not a
"broken" result and I don't want to mark it wrong.
Then, once I have the commit, how do I work out which change within it
caused this?
Automate the test if you can. git bisect run with a script that exits non-zero on failure turns twenty manual checks into one command.
This only happens in production: <describe>
I can't reproduce it locally. Stack: <stack>. Host: <host>.
I don't want to add print statements to a live site. Tell me:
- What to log, temporarily, that would identify this — and how to
scope it to just the affected case rather than every request.
- What existing signals I might already have: access logs, error
reports, host metrics.
- Whether I can capture the full context of one failing request
without recording everyone's data.
- What differs between production and local that could cause this
specifically.
Nothing that logs personal data or secrets.
Log an identifier and the shape of the data, never the contents. You almost always need to know that a field was empty, not what was in it.
A checklist for the differences that actually matter.
dev-vs-prod.txt
Works in development, broken in production.
Symptom: <describe>. Production error: <paste from host logs>.
Stack: <stack>. Host: <host>.
Go through what differs and tell me which could produce this exact
error:
- Environment variables set locally but not on the host.
- File path case sensitivity.
- Working directory and relative paths.
- Runtime version.
- Build step outputs missing or stale.
- HTTPS, secure cookies, mixed content.
- Database version, or a different database entirely.
- Timeouts and memory limits the host imposes.
- Network egress rules.
For each plausible one: exactly how do I check it on the host?
Get the production error first. Everything before that is guessing, and the host's log has the line number your browser doesn't.
My <stack> site shows a blank white page or an error 500 on <host>, and
works fine locally. The live URL: <url>
Help me find the real error:
- Where this host writes the error I can't see.
- How to turn on detailed errors briefly, for me only, without showing
them to visitors.
- The usual causes for this stack on this host: a missing extension,
a different runtime version, file permissions, a path that differs,
a memory limit.
Once I paste the actual error, tell me the fix.
A blank page means an error was hidden, not that nothing happened. Find the log before you change any code.
Can't write the file, can't read the folder, can't run the script.
permission-denied.txt
My <stack> app on <host> fails with a permission error:
<paste the error>
Explain for this host:
- Which user my app actually runs as, and which user owns the files.
- The permissions this path needs, and the smallest change that works.
- Why "make it 777" is the wrong fix, and what to do instead.
- Whether a folder the app writes to should be inside the web root at all.
Then give me the exact commands, or control panel steps, to fix it.
Wide-open permissions make the error go away and leave a hole open. On shared hosting, other accounts on the same server can sometimes use it.
My <stack> site is quick on my machine and slow on <host>. The slow part:
<which pages or actions>
Help me find out why:
- How to measure where the time goes on the host, not locally.
- The usual differences: database on another server, no caching, a
slower disk, a cold start, a small memory limit.
- Queries that are fast on a tiny local database and slow on real data.
- External calls my code makes that are slower from the host's network.
Tell me what to measure first and what each result would mean.
Your local database has fifty rows and the live one has fifty thousand. A missing index doesn't show up until the data does.
Part of my <stack> app takes a while: <describe it>. On <host> it gets
cut off with <the error or behaviour you see>.
Explain:
- Which timeouts are in play here: the web server, the language runtime,
a proxy, the browser.
- Which of them I can raise on this host, and which I can't.
- Whether raising them is even the right fix.
- How to restructure the work so it finishes in the background and the
page checks back, using only what this host supports.
Recommend one approach.
Raising a timeout just moves the wall. If the work can take longer than a person will wait, it belongs in the background.
Sent, according to the code. Nowhere, according to the inbox.
mail-not-arriving.txt
My <stack> app on <host> sends <what kind of email>. People say they
never arrive, or they land in spam.
Help me work out where they're going:
- Is mail leaving the server at all? How do I check on this host?
- Is the host blocking or rate-limiting outgoing mail?
- Are SPF, DKIM and DMARC set up for my domain? What should they say?
- Would a transactional email service be more reliable than sending from
this host, and how would I switch?
Give me the checks in order, cheapest first.
Mail sent straight from a shared host's server often lands in spam, however good your code is. A sending service with the right DNS records usually fixes it outright.
The bug that only exists because the two machines disagree.
version-mismatch.txt
My project uses <stack>. On my machine it works; on <host> I get:
<paste the error>
Check whether this is a version difference:
- How to see the exact runtime and extension versions on both machines.
- Which features in my code need a newer version than the host might have.
- Extensions my code needs that the host may not have enabled.
- How to pin the version I develop against so this can't happen again.
Then tell me whether to change my code or change the host's setting.
Develop on the version the host runs, not the newest one. Being told a feature doesn't exist is a much better way to find out than a blank page in production.