A headless browser that visits a URL and captures an image is the basis of link previews, visual regression testing, and PDF generation. It's also the project where "it works locally" and "it works on the server" differ most.

1. Headless browser, not a screenshot API

Running it yourself means no per-image cost and no third party seeing your pages. It also means dealing with fonts, timing and memory — which is the actual work.

screenshot-service.txt
Build a screenshot service. Headless browser, <stack>.

- Takes a URL, viewport size, and full-page or viewport-only.
- Waits for the page to actually be ready — not a fixed sleep. Explain
  what you're waiting for and why that's more reliable.
- Blocks cookie banners and common overlays where possible.
- Timeout per capture, and a hard cap on concurrent browsers.
- Reuses a browser instance across captures rather than launching one
  each time, but recycles it periodically — they leak memory.

Then: what fonts will be missing on a Linux server, and how do I fix
that? Emoji too.

2. Waiting is the hard part

A fixed sleep 2 is unreliable in both directions — too slow for simple pages, too fast for heavy ones. Wait for the network to go quiet, for specific elements to appear, or for fonts to finish loading. Combine them with a timeout.

Most bad screenshots are timing: captured mid-render, before images loaded, or during an animation.

3. Fonts, on a server

The single most common surprise. Your development machine has hundreds of fonts; a minimal Linux container has almost none. Screenshots come back with fallback fonts, wrong metrics and missing emoji.

Install a font package explicitly, including an emoji font, and verify with a test page containing everything you care about.

4. Treat every URL as hostile

If users supply the URL, you're making your server fetch arbitrary addresses. Block private IP ranges, localhost and cloud metadata endpoints — otherwise your service is a tool for probing your own internal network.

Cap the page size, the timeout and the number of redirects.

5. Memory

Headless browsers are heavy and they leak. Limit concurrency, recycle instances after a number of captures, and set memory limits. Unbounded, this will exhaust a small server within an hour.

Compare screenshots of the same page taken twice before trusting any visual diffing. If two runs of an unchanged page differ, you have a timing problem, not a regression.