Stop it drafting twice
At xhigh or max, a long deliverable can be written in thinking and again in the reply.
Everything you produce in one reply, including any reasoning or drafting before the reply, counts toward a single output limit of about <max_tokens> tokens. Composing the entire deliverable in full as reasoning and then again as the reply would double the length of the turn without improving the result — so don't. Use the reasoning space to understand the request, check the inputs, and settle the structure and difficult decisions. Use the output space to write the thing once.
Append this to the end of the request, with the real token limit filled in. Or simply run long-deliverable requests at high effort, where the problem doesn't arise.