Snippet · Building with models
Handle content you didn't write
User content, model output, or both.
My app displays <user-submitted content / model output> publicly. Design the handling: - What gets checked, before or after publishing, and by what. - Where a provider's moderation endpoint helps and where it doesn't. - What happens to flagged content — blocked, held for review, or published with a warning? - How a false positive gets appealed, since there will be some. - What I log, and for how long. Then, separately: treat this content as untrusted input. Where does it reach the page, a query, or a tool call, and is it escaped at each?
Moderation and escaping are different problems. Content can be entirely inoffensive and still contain a script tag.