In late 2023, we were three people building an internal AI assistant for a workflow that touched regulatory reporting. The assistant was supposed to help analysts prepare structured summaries faster. It was good at that. It was also, periodically, wrong in ways that our compliance team would catch at the end of each sprint review. By the third sprint, we noticed the pattern: the assistant would occasionally include specific data references that fell outside the disclosure scope, or phrase a finding in a way that implied a regulatory conclusion the analyst had not actually drawn.
The compliance review process was doing its job. The problem was that it was doing its job weeks after the responses were generated and in a context where fixing them meant re-running the entire workflow. The review was the wrong place to enforce policy. Policy was being enforced downstream, at high cost, after the fact. We wanted to move it upstream, to the moment of generation.
What We Tried First
The first approach was system prompt engineering. We added detailed behavioral instructions to the model's system prompt: avoid implying conclusions, qualify any regulatory language, do not surface data outside these specific fields. This worked for a while. Then we updated the prompt for an unrelated reason, a tone adjustment, and the qualifier instructions drifted. Two weeks later, compliance caught a response that had reverted to the old behavior. The prompt update had accidentally weakened the behavioral constraints.
The second approach was output filtering with a second model call: take the first model's response, pass it to another model with instructions to check for compliance issues, and only deliver it to the user if the second model approved. This worked technically. It also doubled the latency of every request, and the second model's evaluation was inconsistent: the same response would pass on one call and fail on another depending on the model's sampling state. We were trading one problem for two smaller but harder-to-measure problems.
By January 2024, we had a working internal tool that was slower and less reliable than it would have been without the compliance layer. That was not the outcome we were aiming for. But we had also learned something specific: the problem was not that compliance checking was hard. The problem was that we were using the wrong tool for the job. Prompt engineering and model-based checking are good for behavioral guidance; they are not good for deterministic rule enforcement. A rule is either applied or it isn't. You need something that behaves deterministically, not probabilistically.
The Insight That Changed the Architecture
Priya had worked on API gateway systems at her previous role, specifically on rate limiting and content routing at the infrastructure layer. The insight she brought to this problem was: policy enforcement for AI responses is structurally the same problem as content routing in a gateway. You have a request, you have a response, you have a set of rules about what the response is permitted to contain, and you need to check those rules synchronously before delivery.
The difference from a standard API gateway is that the "content" is natural language rather than structured data. That changes how you write the matching rules. But the evaluation architecture, synchronous intercept at the response layer, rule evaluation against active policies, action-on-match, is the same pattern. We built ZeroDrift as an inference-layer proxy that applies this pattern to LLM outputs.
Marcus brought the compliance engineering background. He had spent years in financial services compliance automation and had a precise sense of what policy rule definitions needed to look like to be legally defensible. Not just "does this response contain a prohibited term," but "does this response make a specific type of assertion in a specific context that falls outside the permitted scope." That nuance mattered for the kinds of deployments we were targeting: regulated industries where "the model might say something bad" is not specific enough to satisfy a legal or compliance team reviewing the deployment.
What ZeroDrift Is Actually For
We built ZeroDrift for a specific problem: teams shipping AI assistants into workflows where the responses have real-world consequences, and where "we reviewed the model's behavior during development" is not sufficient assurance for the people responsible for the outcome. That is a real and growing category of deployment. We are not building a safety tool for general AI usage. We are building compliance infrastructure for teams that already know what their AI assistant is supposed to do and need a reliable way to enforce those boundaries at runtime.
The product we have built is a middleware layer that sits between your application and your inference provider. When a response comes back from the model, ZeroDrift evaluates it against your active policy ruleset before it reaches the user. Responses that violate policy are either blocked or rewritten, depending on the rule action. Everything is logged with a tamper-proof audit trail. The latency overhead at p99 is under 30 milliseconds for evaluation-only events.
We are not claiming that ZeroDrift solves every AI risk problem. It solves the enforcement gap: the space between "we defined what the assistant should not do" and "the assistant actually doesn't do it, every time, verifiably." That gap is where most enterprise AI deployments currently live. Closing it is not glamorous work. It is necessary work, and it is what we set out to do.
What We Have Learned Building This
Eighteen months of building and running early access deployments have taught us that the compliance problem is not primarily a model problem. The models are capable enough. The gap is organizational: the people who define policy and the people who build the AI product are not, in most teams, working from the same enforcement layer. The compliance team writes a policy document. The engineering team writes a system prompt that tries to reflect it. The gap between those two representations is where violations happen.
ZeroDrift does not require the engineering team to translate policy into a system prompt. Policy is defined in the policy layer, in terms that a compliance team can review and own independently. Engineering writes code that calls our API. The two concerns are separated at the architecture level, not just in the documentation. That separation is, in our view, the actual requirement for sustainable AI compliance in regulated deployments.
We are a small team. We have not raised outside capital. We are building this product because we needed it and could not find it. If you are building an AI assistant for a regulated workflow and you have encountered the compliance gap, we want to talk.