Sorting things automatically
Classification is where AI is most reliable and least glamorous.
Build time ~4 hrs
Classification — routing tickets, tagging content, flagging what needs a human — is the most dependable thing to build with a model. The output space is small, the answers are checkable, and you can measure whether it's working.
1. Define the categories properly
Most classifier failures are category failures, not model failures. Categories must be mutually exclusive, collectively exhaustive, and distinguishable by someone reading only the input.
Include an explicit "unclear" or "other" option. Without one, ambiguous items get forced into a category and you never find out.
2. Write the labelled set first
Fifty examples you've labelled by hand. This is the boring step everyone skips, and it's the only thing that tells you whether your classifier works.
You'll also discover your categories are wrong while doing it, which is much cheaper to find out now.
Classify the input into exactly one of these categories:
<category>: <precise definition, and what it excludes>
<category>: <precise definition, and what it excludes>
unclear: use this when the input genuinely fits none of the above, or
fits more than one equally.
Return JSON: {"category": "...", "confidence": "high|medium|low",
"reason": "one short sentence"}
Do not invent categories. Prefer "unclear" over a forced fit.
Input: <text>3. Measure against your set
Run all fifty. Count correct. Look at every mistake — the pattern in the errors tells you whether to fix a category definition, add an example to the prompt, or accept the limit.
Re-run this whenever you change the prompt. It takes a minute and catches the regressions that "it seems better" hides.
4. Route the uncertain ones to a human
The design that makes this safe. High confidence gets acted on; low confidence and "unclear" go to a queue. You get most of the automation with none of the confident-wrong failures, and the queue tells you where to improve.
5. Keep the reason
Storing the one-line reason alongside every classification costs almost nothing and makes the system debuggable. When something is filed wrongly, you can see what it was thinking.
Track what proportion goes to the human queue over time. If it climbs, your inputs have changed and your categories need revisiting.