There are two things you can do when a compliance check identifies a policy violation in an LLM response. You can block it, returning either an error or a standardized fallback message to the user. Or you can rewrite it, replacing the violating content with compliant content and delivering the modified response as if nothing unusual happened. Both approaches protect policy. They do so with substantially different tradeoffs, and the right choice varies by violation category, user context, and how much policy engineering investment you have available.
This piece is a direct comparison of the two approaches, where each performs well, where each fails, and how production systems typically combine them.
The Case for Filtering
Filtering is operationally simple and has clear failure modes. When a violation is detected, the violating response is discarded. The user receives a message that the assistant cannot help with that query, or a redirect to an appropriate resource. The compliance record shows the blocked response and the triggering rule. Everyone knows what happened.
The strongest argument for filtering is certainty. A blocked response cannot contain a violation because it is not delivered. A rewritten response may contain a violation if the rewrite logic is wrong, incomplete, or was applied to a violation the rewrite template was not designed to handle. For violation categories with severe consequences (regulatory exposure, data privacy obligations under CCPA or HIPAA, specific financial advice prohibitions), the certainty of blocking often outweighs the conversational cost.
Filtering is also the right answer when the underlying query should not be answered at all. An AI assistant deployed in a benefits administration context should not attempt to answer detailed tax questions, even with heavy qualifications. The correct response is to acknowledge the question and redirect. Rewriting a tax-advice response into a compliant version that still attempts to address the tax question is threading a needle that does not need to be threaded.
The costs are real, though. Frequent blocking visibly degrades the user experience. Users who see "I can't help with that" responses repeatedly will either lose trust in the assistant or learn to rephrase queries to evade the filter, neither of which serves the deployment's goals. For violation categories that come up regularly in legitimate queries, blocking is essentially capping the assistant's usefulness in that category.
The Case for Rewriting
Rewriting preserves the conversational value of the response while removing the specific element that constitutes a violation. Done well, the user receives a useful answer; the compliance record shows the violation that was caught and the specific transformation applied; and the interaction continues naturally. For many violation categories, this is a meaningfully better outcome than blocking.
The violation categories where rewriting is most valuable are those where the underlying information is legitimate but requires specific treatment to be compliant. A response that provides accurate healthcare information without the required disclaimer can be rewritten to add the disclaimer without changing the information content. A response that mentions a competitor by name can have the mention replaced with a category reference. A response that quotes specific account balances can have the balance figures redacted while preserving the rest of the answer. In each case, the rewrite serves both the user and the compliance requirement.
The quality of a rewrite depends heavily on the specificity of the transformation. Template-based rewrites, where the policy team has defined a specific substitution for a specific violation pattern, produce consistent and predictable output and run fast enough for synchronous inline use. Generative rewrites, where a second model produces a revised response, can handle novel violation patterns and produce more natural-sounding output, but introduce their own quality risks: the generated rewrite may still contain a violation in different language, may lose important information from the original response, or may introduce new problems not present in the original.
The Hidden Risk in Rewriting
Rewriting introduces a failure mode that filtering does not: the rewrite itself can be wrong. This is not a theoretical concern. It shows up in practice in two main ways.
The first is incomplete transformation. A rewrite template that addresses a specific violation pattern may not address all instances of that violation in a given response. A response that contains multiple PII disclosures at different structural locations in the text may have only the first one caught by the detection logic and rewritten, with subsequent ones passing through. Rewriting a violation in sentence three does not mean sentences six and nine are clean.
The second is transformation artifacts. A rewrite that replaces specific language with approved language can create grammatically awkward or logically inconsistent sentences if the substitution is not carefully crafted. A response rewritten to remove a balance figure but that still contains references to "the balance shown above" is now internally inconsistent. Users notice this, and it creates a different kind of trust problem.
Neither of these means rewriting is worse than filtering. They mean that a rewrite pipeline needs its own quality assurance layer: a validation pass after rewriting to confirm that the known violation was addressed and that the rewrite did not introduce new issues. This adds complexity and latency, but it is the only way to have confidence that the delivered response is actually clean.
How Production Systems Typically Combine Both
The most common production configuration in ZeroDrift deployments uses a tiered decision model: violation severity and confidence determine the action, not a single global policy.
High-severity violations with clear detection signals (PII disclosure patterns, explicit prohibited terms, direct financial advice with specific figures) map to block actions for severity tiers that represent regulatory exposure, and to template-based rewrites for severity tiers where the information is legitimate but needs qualification. Low-severity violations with lower detection confidence map to rewrites with validation passes, accepting the complexity cost in exchange for preserving response quality.
The coverage question is important: for which violation categories does the policy team have the bandwidth to build and maintain rewrite templates? Template-based rewrites require ongoing maintenance. When the policy changes, the template may need to change. When the violation pattern shifts (because the model is updated, or user query patterns evolve), the detection logic and the template may need to change together. This is real ongoing work, and teams that underestimate it find their rewrite coverage decaying over time.
One practical heuristic: start with blocking for all categories, ship, and measure which blocked categories are producing a high volume of legitimate queries that users actually need answered. Those are the candidates for rewrite investment. This avoids building rewrite infrastructure for categories that rarely trigger or that users are not actually trying to address. Build to the actual failure modes you observe, not the failure modes you anticipate.
What Each Approach Signals to Auditors
Both filtering and rewriting produce compliance records, but the records have different shapes and tell different stories to a compliance reviewer.
A log showing only block events is easy to interpret: violation detected, response discarded, user received fallback. The audit question is whether block rates are appropriate and whether the blocking logic is correctly identifying violations vs. producing false positives. High block rates on a category that should rarely violate suggest a detection problem. Low block rates on a category that frequently appears in user queries suggest a detection gap.
A log showing rewrite events is richer and more complex. The reviewer wants to see: the original violating response, the specific rule triggered, the transformation applied, and the delivered response. They want to verify that the transformation was complete and that the delivered response does not contain the violation. This is more information to audit, but it is also a more complete demonstration that the system is handling violations thoughtfully rather than simply blocking them.
Neither approach is inherently more favorable from an audit perspective. The important thing is that the records are complete, tamper-proof, and navigable. A compliance reviewer who can pull up any flagged event and see the full chain of detection, decision, and action will have a substantially easier time than one working from incomplete logs, regardless of whether the predominant action was block or rewrite.