Snippet · Building with models
Chunk documents properly
Most retrieval failures are chunking failures.
My documents: <format>, roughly <n> docs averaging <length>. Write the chunking step: - Split on structure — headings, sections, paragraphs — not a fixed character count that cuts sentences in half. - Target <n> tokens with sensible overlap, so a fact spanning a boundary isn't lost. - Keep with each chunk: source, heading path, position in document. - Handle documents shorter than one chunk, and single sections longer than a chunk. Then give me a way to print 20 random chunks so I can read them. I want to check each would make sense to someone seeing it in isolation.
Read the chunks. It's the step everyone skips and the one that finds the problem — usually that headings were stripped and half the chunks are context-free fragments.