"Ask questions of my documents" is the most requested AI feature and the one most often built backwards. The model is not the hard part. Finding the right passage to give it is the hard part, and it's where nearly every failure originates.

1. Chunk with care

Documents get split into passages before anything else happens. Split badly and no amount of prompting recovers.

Split on structure — headings, sections, paragraphs — rather than on a character count that cuts sentences in half. Keep chunks in the range of a few hundred words. Overlap them slightly so a fact spanning a boundary isn't lost. And keep the source and heading with every chunk, because you need them for citations.

chunking.txt
My documents are: <format — Markdown / PDFs / HTML>, roughly <n>
documents averaging <length>.

Write the chunking step:
- Split on document structure, not a fixed character count.
- Target <n> tokens per chunk with sensible overlap.
- Keep with each chunk: source file, heading path, and position.
- Handle documents shorter than one chunk, and single sections longer
  than a chunk.

Then show me how to eyeball the output — I want to read 20 random
chunks and see whether they'd make sense to a stranger in isolation.

That last step is the one everyone skips and it's the one that finds the problems.

2. Retrieve, then check retrieval on its own

Embed the chunks, store them, search by similarity. But before connecting a model at all: type twenty real questions and look at what comes back.

If the right passage isn't in the top few results, the answer will be wrong no matter what you do afterwards. Fix retrieval first. Better chunking, or combining keyword search with vector search, usually helps more than anything else.

3. Then generate, with citations

Give the model the retrieved chunks and the question, and instruct it plainly: answer only from the provided passages, cite which one each claim came from, and say "I don't know" when the passages don't contain the answer.

Show the citations in your interface. It's what makes the feature trustworthy, and it lets users check you.

4. Measure it

Write twenty questions with known correct answers. Run them after every change. Without this you're tuning by vibes, and every improvement to one question quietly breaks another.

When an answer is wrong, always check the retrieved chunks first. Nine times out of ten the model was reasoning correctly about the wrong material.