Adding a model to something costs money per use, adds latency, introduces non-determinism, and creates a failure mode where output is confidently wrong. It's worth it where those costs buy something genuinely difficult otherwise.

Where it earns its place

Unstructured to structured. Messy text in, clean data out. Parsing ingredients, extracting fields from emails, categorising tickets. Reliable, checkable, hard to do any other way.

Classification and routing. Small output space, verifiable answers, measurable accuracy. The most dependable category.

Summarising for a specific reader with a specific decision. Not generic summaries — those help nobody.

Search over meaning rather than keywords. Finding the relevant document when the user's words don't match the document's words.

Removing a blank page. A first draft the user edits. The bar is low because they're going to change it anyway.

Where it usually isn't

Things a rule handles. If the logic is expressible in code, code is cheaper, faster, deterministic and free.

Anything requiring exactness. Arithmetic, lookups, anything where a plausible wrong answer causes harm. Use the model to decide what to look up, then look it up properly.

Chat as an interface to a thing with three buttons. Conversation is a slower way to do something a form does well.

Features added because the technology is available rather than because a user needs the outcome.

The test

ai-feature-test.txt
I'm considering adding this AI feature: <describe>

Answer honestly:
- Could a rule, a lookup or a search do this? If so, why is a model
  better here specifically?
- What does a wrong answer cost the user, and would they notice it
  was wrong?
- What's the cost per use at my expected volume?
- What's the experience when the model is slow or unavailable?
- Would users choose this over a well-designed form?

If this is a bad idea, say so plainly.

Design for wrongness

Everything a model generates should be reviewable, correctable and attributable. Show sources. Let people edit output. Make it obvious what was generated. Route low-confidence cases to a human.

That design — assistive, not autonomous — is what makes AI features trustworthy rather than a liability.

If the feature only works when the model is right, and there's no way for the user to tell whether it was, it isn't finished.