A compliance rewrite that replaces a useful answer with a generic deflection has not solved the compliance problem. It has traded one problem for another. The original response was violating policy. The rewritten response is policy-safe but now useless, and the user who asked a legitimate question is worse off than if the assistant had not tried to answer at all.
This failure mode is common in early policy engineering work. The goal becomes "make this response pass the rule check" rather than "deliver the policy-safe version of what the user actually needed." Those two goals produce very different rewrite templates. Getting from the first to the second requires thinking carefully about what the violation was, what the permissible version of the answer looks like, and where those two things differ.
Violation Extent Is More Specific Than It Looks
Most rewrite mistakes happen because the rule writer treated the entire response as violating when only a specific extent of it did. Consider a response that answers a user question about investment strategies in general terms, then closes with a sentence that could be read as specific financial advice. The violating portion is that last sentence. The rest of the response is fine.
A well-designed rewrite rule identifies the violation extent: the specific span of text that breaks policy. The replacement targets only that extent. The result is a response that still answers the question, still provides useful context, and terminates with the policy-compliant version of the closing: a disclaimer, a referral, or a rephrased general statement that cannot be read as specific advice.
Over-broad rewrites, which replace the entire response, are a sign that the rule writer did not clearly define the violation extent. The rewrite template works around a match condition that is too coarse. The fix is a tighter match: a more specific pattern that targets the actual problematic construct, not everything that loosely resembles it.
Three Violation Categories and Their Rewrite Approaches
Violation types have natural rewrite strategies. Getting these wrong inverts the benefit of rewriting over blocking.
The first category is unsolicited assertion: the model states something the company cannot stand behind, typically a factual claim in a regulated domain. The rewrite approach is not to remove the answer. It is to hedge the assertion: replace "this product has X characteristic" with "this product is generally characterized as X; please consult the product documentation for current specifications." The useful information stays. The unqualified assertion is gone.
The second category is topic bleed: the assistant wanders into a prohibited subject area mid-response. The user asked something adjacent, and the model's answer drifted into off-limits territory. The rewrite approach here is a truncation plus redirect. The portion up to the drift point is often entirely usable. The drift is replaced with a closing that acknowledges the user's question and redirects appropriately. The redirect should be specific enough to be useful: not "I cannot help with that" but "For questions about X, I can point you to [resource]."
The third category is missing qualification: the response is substantively correct but lacks a required element, such as a disclaimer, a caveat about professional advice, or a consent disclosure. This is the cleanest rewrite case. The original response content is preserved entirely. The required qualification is appended or inserted at the appropriate point. The rewrite template is deterministic and the compliance improvement is precise.
Writing Replacement Templates That Actually Read Well
The usability of a rewrite depends heavily on how the replacement text was written. Replacement text drafted by a compliance team without editorial review often reads like a legal notice inserted into a conversation. Users notice the register shift. It damages the assistant's perceived quality even when the answer is technically correct.
The fix is simple but requires an intentional step: write your replacement templates in the same register as the assistant's normal output. If the assistant uses a conversational tone, the disclaimer should use a conversational tone. "This information is provided for general guidance only and does not constitute financial advice; you should consult a qualified advisor for decisions relevant to your specific situation" says the same thing as a formal legal notice, but reads as something a helpful colleague might actually say.
When reviewing rewrite outputs in our early access program, the quality signal we used was: "could this sentence have appeared in the original model response, or does it read like an insertion from a different document?" If it reads like an insertion, the template needs rewriting, not the rule.
When Rewrite Is the Wrong Tool
Not every violation is a good candidate for rewrite. Some violation types have no useful policy-safe version to produce. These are the cases where block is the right action, not rewrite.
The clearest example: a request that is itself a policy violation, where any substantive response would require engaging with the prohibited topic. Rewrites for these cases tend to produce either a generic deflection that adds nothing, or they accidentally preserve more context from the original response than intended. In these cases, block-with-redirect is better. The block is an explicit outcome. The redirect tells the user where to go instead.
The selection between rewrite and block should be an explicit design decision at the rule level, not an afterthought. For each violation type in your policy, ask: is there a genuinely useful policy-compliant version of this response? If the answer is yes, design the rewrite template. If the answer is no, or if the useful version is so stripped down that it adds nothing, design a block-with-redirect instead.
Policy Review for Rewrite Quality, Not Just Compliance
After your policy ruleset has been running for a few weeks, the most valuable review is not of violations caught. It is of rewrites delivered. Pull a sample of 50 responses where a rewrite was applied and read them as a user would. Ask: did the user get a useful answer? Does the response read naturally? Is the rewritten segment clearly worse quality than the surrounding text?
Rewrites that degrade quality consistently are telling you something about the template, not the model. The model produced a response. The template replaced part of it poorly. This is fixable at the policy layer without touching the model or the system prompt.
The goal of a good rewrite policy is that a user, reading the final response, cannot tell that a compliance layer was involved. The answer should read as if the assistant simply answered correctly the first time. That is a high bar, and hitting it requires iteration on templates with real output samples. But it is the correct bar: compliance enforcement that is invisible to users is enforcement that does not damage the product it is protecting.