Snippet · Building with models
Stream model responses through my host
Some hosts buffer the whole response, so streaming arrives all at once.
I'm streaming responses from <model> through my <stack> app on <host>. Locally the text arrives word by word; live it arrives all at once. Help me find what's buffering it: - Output buffering in the language runtime. - Compression on the web server, which often waits for the whole response. - A proxy or CDN in front of the host. - Headers that tell each layer not to buffer. Give me the settings for this host. If it can't stream at all, show me a fallback that polls for progress instead.
Compression is the usual culprit. The server waits to compress the whole response before sending any of it.