Snippet · Building with models
Guardrails on a model in production
It's talking to strangers on your behalf.
My app exposes a model to users here: <describe the feature> Add the practical protections: - Input length cap, and a rate limit per user and per IP. - A hard spend cap across all users that fails closed, with an alert before it's reached. - Treat model output as untrusted: never render it as HTML, never execute it, never put it into a query. - If user content is passed into the prompt, structure it so instructions inside that content aren't followed. Show me how. - Log the prompt and response for anything that errors or gets flagged, with personal data excluded. Tell me what someone would try first to abuse this, and which protection stops it.
Model output is user input from somewhere else. If you render it as HTML, you've built a cross-site scripting hole with extra steps.