MEMON SYSTEMS

Generation-Tier Integrity in Legal RAG: An In-Flight Verification Control Specification

Profile Commercial Legal-RAG Configuration — External Architecture Analysis (Italy)
Stack Specified control: Llama-3-70B on vLLM, RadixAttention, 8B NLI auditor, ColBERTv2, AST semantic splitting. Analysed system: reconstructed from public documentation
Duration 3 Weeks — Architecture Analysis & Design
By design
0% contaminated output and a per-query generation-tier audit trail, following from selective KV-cache truncation where the cache is exposed. Nothing here was measured: the control has not been built or deployed, and the cost and latency figures are analytical models over published pricing.

The Challenge

A commercially marketed legal-RAG platform serving 6M+ Italian and EU legal sources advertises grounded output to practising lawyers. Reconstructed from its public product and engineering documentation, the generation path uses a post-hoc verification gate that evaluates the complete response after all tokens are generated. When an LLM produces an ungrounded token sequence mid-response — a structurally predictable event occurring on 17–33% of specialised legal queries (Stanford RegLab) — the contaminated tokens persist in the generation cache and condition every subsequent token. The gate catches the contamination after propagation, forcing a full regeneration cycle at full API billing cost. The result is an uncontrolled OPEX penalty on the generation path — a compliance control deficiency invisible to aggregate monitoring but compounding linearly with query volume.

Important

Provenance & Status. This is a reference architecture and deterministic cost model, not a deployment report. Its evidentiary basis is stated here in full so nothing below has to be taken on trust.

No vendor engagement, no privileged access. I was not retained by, and had no access to, the platform analysed here. The architecture in §3 is reconstructed from public sources — the vendor's published product description, public engineering requirements, and the framework defaults those imply. §4.1 classifies every input as Confirmed or Inferred with its reasoning chain. The vendor is not named, and nothing here derives from confidential information.

The remediation is specified, not deployed. The in-flight auditor in §6 is a reference architecture I authored. It has not been built. No GPU cluster was provisioned, no shadow migration was run, and no production traffic has passed through it. §11 is a rollout protocol with acceptance gates, not a rollout report.

Every number carries a provenance label. Four classes are used throughout, and no figure of one class is presented as another:

Label Meaning
Confirmed Sourced from published vendor documentation or published third-party research.
Inferred Derived from Confirmed facts by a stated reasoning chain (§4.1).
Structural True by construction — provable from the architecture's semantics without running traffic.
Modeled Arithmetic from stated inputs (published pricing, declared token assumptions).

What is actually proven here. The load-bearing claim is Structural and needs no benchmark: behind a managed API boundary, partial remediation of a contaminated generation is impossible, so the cost floor per hallucination event is a full regeneration (§12.1). Everything downstream — the 95% waste reduction, the 92.2% cost reduction, the 120ms rollback — is Modeled from that structural fact plus published pricing. §15 states what would falsify the model.


1. Executive Summary

A commercially marketed legal-tech platform operating a Conversational RAG service over 6M+ Italian and European legal sources ships fabricated legal text to practising lawyers on a structurally predictable basis. Its generation path relies on a post-hoc verification gate — the industry-standard "generate-then-validate" pattern — which evaluates the full LLM response after generation completes. When the model produces a hallucinated citation or ungrounded legal holding mid-response, the contaminated tokens persist in the generation state and condition every subsequent token. The gate detects the contamination only after the damage has propagated through the entire output.

Note

How that architecture was determined. I did not observe this system's internals. The post-hoc pattern is Inferred (High) from two Confirmed facts: the platform's published dependency on proprietary frontier APIs, and its published use of standard orchestration frameworks whose default validation pattern is generate-then-validate. The full reasoning chain, with each link classified, is in §4.1. This is an architectural reconstruction from public evidence — a well-supported one, but a reconstruction. Where a reader has better information about a specific vendor, the specific claim should yield; the structural argument in §12.1 does not depend on it.

Under the SRA's Principles, a solicitor relying on AI-generated legal citations bears professional liability for their accuracy — regardless of whether the error originated in the model or the retrieval tier (Principle 2: uphold public trust; Principle 7: act in the client's best interests). The Italian Codice Deontologico Forense imposes analogous obligations under Articles 9 and 12, holding the avvocato personally liable for the accuracy of materials presented to a tribunal. Peer-reviewed research from Stanford's Regulation, Evaluation, and Governance Lab (Magesh et al., 2025) documents hallucination rates of 17–33% in specialised legal RAG systems — meaning one in five to one in three queries structurally triggers this failure. (Confirmed — published, peer-reviewed, domain-specific.) Under the Mata v. Avianca precedent, US federal courts have imposed per-citation sanctions of $1,000–$5,000 for fabricated case references. For a legal-tech platform marketing to enterprise law firms, each uncontrolled hallucination event that reaches an end user is not a performance degradation — it is a sanctionable compliance incident with compounding professional liability exposure.

This document specifies an in-flight generation-tier compliance control — an inference-time sentence-level auditor with selective KV cache truncation — that eliminates contaminated token sequences from the generation state before downstream tokens are conditioned on them. The control is a design. It is not running, not deployed, and not monitored. What is established is the architectural case for it: the mechanism is possible with an exposed KV cache and provably impossible without one (§12.1).

The generalised principle: Post-hoc verification is a compliance control that fires after the damage has propagated. In regulated domains where output accuracy carries professional liability, the verification boundary must operate during generation — not after it — to prevent contaminated state from reaching the output tier.


2. Regulatory Context

The failure analysed here — uncontrolled hallucinated legal text reaching practising lawyers via an AI-assisted research platform — intersects multiple regulatory frameworks that govern AI system accuracy, human oversight, and professional liability in regulated domains.

Regulatory Framework Relevant Provision How This Failure Class Violates It
EU AI Act Article 15 — Accuracy, robustness, cybersecurity System produces fabricated legal citations on a structurally predictable query class with no mechanism to prevent contaminated output from reaching the user
EU AI Act Article 14 — Human oversight No inference-time control exists to intercept or flag ungrounded generation before output delivery; the post-hoc gate either ships the contamination or imposes a full regeneration with no user visibility
EU AI Act Article 9 — Risk management system Hallucination guarantee not validated against structured adversarial conditions; the 17–33% baseline rate (Stanford) is a known, documented risk with no mitigation control on the generation path
Italian Codice Deontologico Forense Articles 9 & 12 — Diligence and accuracy obligations An avvocato relying on fabricated AI-generated holdings bears personal disciplinary liability; the platform provides no mechanism to distinguish verified from unverified output
SRA (UK) Principle 2 — Uphold public trust AI system presents hallucinated legal text as authoritative citation with full confidence and zero disclaimer
SRA (UK) Principle 7 — Act in the client's best interests Practitioner receives fabricated legal holdings that could directly harm client interests if relied upon
ABA (US) Model Rule 1.1 — Competence Attorney relying on system receives contaminated legal analysis; duty of competence extends to verifying AI-assisted research tools
NIST AI RMF GOVERN 1.1 — Risk management processes No documented control for the identified generation-tier hallucination risk; no structured audit trail for verification decisions
ISO 42001 A.6.2.6 — AI quality objectives No measurable generation accuracy objective defined, monitored, or enforced on the inference path

Warning

The regulatory exposure is bilateral. The platform operator faces liability under the EU AI Act for deploying a high-risk AI system without adequate accuracy controls (Article 15). The practising lawyer faces professional disciplinary liability under their respective bar association for relying on unverified AI output. A single documented instance of fabricated legal text reaching a tribunal creates exposure for both parties simultaneously.

Note

This table is a regulatory analysis of the failure class, not a legal opinion on any specific vendor's compliance status, and not the output of a formal conformity assessment. I am an architect, not counsel.


3. The Failure — What Is Happening

3.1 What the User Experiences

A practising lawyer submits a query to the RAG platform requesting analysis of a specific Italian statutory provision. The platform returns a response containing fabricated legal holdings, misattributed citations, or ungrounded interpretive commentary — presented with full confidence, correct formatting, and zero disclaimer. The lawyer has no mechanism to distinguish verified output from hallucinated output. If the lawyer relies on this response in a brief, filing, or client advisory, they are presenting fabricated law as authoritative — a potentially sanctionable act under every major bar association's professional conduct rules.

In a subset of queries (approximately one in five, per the Stanford baseline rate), the user experiences a secondary failure: the verification gate catches the contamination post-generation and triggers a full regeneration cycle. The user sees a Modeled 12–25 second response delay as the system silently discards the contaminated output, re-processes all input tokens, and regenerates the entire response from scratch. In a conversational interface, this delay is indistinguishable from a frozen application. Users abandon the session, and the wasted compute is billed regardless.

3.2 Why It Is Dangerous

This is not a random model error that can be addressed by scaling to a larger model or improving prompt engineering. The failure is structural — it is a mechanical consequence of how autoregressive language models generate text when combined with post-hoc verification:

  1. KV cache contamination is invisible. Large Language Models are autoregressive: each token is generated conditioned on the full sequence of preceding tokens. The key-value (KV) cache stores the attention states for all previously generated tokens. When a hallucinated token sequence is written mid-response, those attention states persist in the cache and influence every subsequent token. The contamination is not confined to the hallucinated sentence — it structurally degrades all downstream generation.

  2. Post-hoc verification creates a compliance-free zone. The verification gate evaluates the completed response. Between the first hallucinated token and the moment the gate fires, there is no compliance control active on the generation path. Every token generated in this interval is conditioned on contaminated state, amplifying the fabrication.

  3. The trade-off is existential. The platform faces a binary choice with no middle ground:

graph TD
    A["Post-Hoc Verification Gate<br/>fires after full generation"] --> B{Hallucination<br/>Detected?}
    B -->|Yes — Strict Threshold| C["Option A: Catch & Regenerate<br/>━━━━━━━━━━━━━━━━━━━━<br/>Discard all tokens.<br/>Re-process all input tokens.<br/>Re-generate entire response.<br/>━━━━━━━━━━━━━━━━━━━━<br/><b>Latency: 12–25 seconds</b><br/><b>User assumes app is frozen</b><br/><b>Session abandonment</b>"]
    B -->|No — Lowered Threshold| D["Option B: Ship Contamination<br/>━━━━━━━━━━━━━━━━━━━━<br/>Hallucinated tokens pass gate.<br/>Fabricated citations reach user.<br/>Lawyer relies on fake holdings.<br/>━━━━━━━━━━━━━━━━━━━━<br/><b>Professional liability</b><br/><b>Sanctionable event</b><br/><b>TPRM disqualification</b>"]
    B -->|Not detected| D

    style C fill:#593e13,stroke:#000,stroke-width:2px
    style D fill:#4d1414,stroke:#000,stroke-width:2px

Caution

The Verification Trap: Option A kills the product via latency-induced churn. Option B kills the product via accuracy-induced liability. Adjusting the verification threshold cannot resolve this — it merely shifts between two failure modes. The architecture requires a structural intervention, not a tuning adjustment.

3.3 What Is Happening Technically

The platform's generation path follows the standard Retrieve-then-Generate pipeline with a post-hoc verification gate:

flowchart LR
    A[User Query] --> B[Retrieve Chunks]
    B --> C[Generate Full Response]
    C --> D[Post-Hoc Gate]

    D --> E[User]
    D -->|Re-run<br>If Fails Gate| C

The post-hoc gate operates as a binary pass/fail checkpoint — either the entire response is shipped, or the entire response is discarded and regenerated. There is no mechanism for granular, per-sentence remediation.

The critical constraint — and this part is Structural, not inferred: the platform generates responses via proprietary frontier APIs (OpenAI GPT-4o, Anthropic Claude). These managed inference endpoints do not expose KV cache state. The application cannot inspect, modify, or selectively truncate the generation cache during inference. When a hallucination is detected post-generation, the only available remediation is complete regeneration — all input tokens re-processed, all output tokens re-generated from scratch. Partial rollback is architecturally impossible on managed API endpoints. This holds for any application on any managed endpoint, regardless of implementation quality, and it is the fact the entire cost argument rests on.

Furthermore, the verification model itself introduces a reliability ceiling. Cross-encoder NLI models (DeBERTa-class) carry a hard 512-token input limit. Italian legal clauses — with their characteristic nested qualifications, cross-references, and conditional exceptions — routinely exceed this boundary. The auditor either truncates the qualifying language that changes the clause's meaning, or fragments verification across multiple passes and loses cross-reference coherence. At the density of a 6M+ legal source corpus, this is a structural reliability cliff, not a theoretical edge case.


4. The Forensic Evidence — The Hallucination Tax

This section presents the structural economic analysis of the failure. The "Hallucination Tax" is a per-query OPEX penalty — a permanent, compounding cost created by the interaction between autoregressive generation, post-hoc verification, and managed API billing mechanics.

Important

This section is Modeled, from Confirmed and Inferred inputs. It is arithmetic over published pricing and declared token assumptions — not an extract from anyone's billing console. I have no visibility into the vendor's actual spend, traffic, or token distributions. Every input is classified below precisely so the model can be re-run against better numbers.

4.1 Data Classification

Every data point used in this analysis is classified by source. Nothing enters the economic model unless it traces to a confirmed source or a high-confidence inference with a stated reasoning chain.

# Data Point Classification Source
C1 Frontier LLM dependency: OpenAI API, Anthropic Claude, Google Vertex AI Confirmed Published engineering requirements
C2 RAG pipeline on top of frontier models, not training from scratch Confirmed Published product architecture
C3 LangChain + LlamaIndex orchestration Confirmed Published engineering requirements
C4 6M+ Italian/EU legal source corpus Confirmed Published product description
C5 Per-token API billing model Confirmed Follows from C1 — managed APIs bill per token
I1 Post-hoc verification pattern (generate → validate) Inferred (High) C3 — framework defaults to generate-then-validate; no public evidence of custom inference-time control
I2 Full regeneration on hallucination detection Structural C1 — proprietary APIs do not expose KV cache; partial rollback is architecturally impossible, not merely unimplemented
I3 17–33% hallucination rate Confirmed Stanford RegLab (Magesh et al., 2025) — peer-reviewed, domain-specific; measured on comparable systems, not on this one
I4 ~800 output tokens per legal analysis response Assumed (Medium) Legal analysis responses are structurally verbose — citations, qualifications, cross-references
I5 ~5,000 input tokens per query Assumed (Medium) RAG retrieval of 3–5 chunks × 500–800 tokens + system prompt + conversation history
I6 r=20%r = 20\% hallucination rate used in the model Assumed Midpoint of the Confirmed I3 range, applied to this platform by analogy

Note

I1 is the weakest link in the chain, and it is load-bearing for the diagnosis. If this platform has in fact built a custom inference-time control, the diagnosis does not apply to it — though the analysis still applies to the very large number of legal-RAG systems built on framework defaults. I2 is the strongest link, and it is load-bearing for the remediation: it is not an inference about anyone's implementation but a property of the managed API boundary itself.

4.2 Pricing Anchors (Confirmed — Published List Prices)

Model Input ($/1M tokens) Output ($/1M tokens) Source
GPT-4o $2.50 $10.00 Published pricing
GPT-4.1 $2.00 $8.00 Published pricing
Claude Sonnet 5 $3.00 $15.00 Published pricing
Claude Opus 4.8 $5.00 $25.00 Published pricing

List prices exclude enterprise discounts, committed-use agreements, and batch pricing, all of which a platform at scale would likely negotiate. The model is therefore an upper bound on per-query API cost.

4.3 Per-Query Cost Model: The Hallucination Tax

Let CgenC_{\text{gen}} denote the cost of a single generation pass (input processing + output generation), rr the hallucination rate, and CwasteC_{\text{waste}} the cost of wasted output tokens on the discarded response.

Cgen=TinPininput cost+ToutPoutoutput costC_{\text{gen}} = \underbrace{T_{\text{in}} \cdot P_{\text{in}}}_{\text{input cost}} + \underbrace{T_{\text{out}} \cdot P_{\text{out}}}_{\text{output cost}}

For GPT-4o with the stated assumptions (Tin=5,000T_{\text{in}} = 5{,}000, Tout=800T_{\text{out}} = 800):

Cgen=(5,000×$2.50/1M)+(800×$10.00/1M)=$0.0125+$0.0080=$0.0205C_{\text{gen}} = (5{,}000 \times \$2.50/\text{1M}) + (800 \times \$10.00/\text{1M}) = \$0.0125 + \$0.0080 = \$0.0205

The Hallucination Tax is the per-query expected cost of the regeneration cycle plus the wasted output tokens on the discarded response:

Ctax=rCgen+rToutPoutC_{\text{tax}} = r \cdot C_{\text{gen}} + r \cdot T_{\text{out}} \cdot P_{\text{out}}

Ctax=0.20×$0.0205+0.20×800×$10.00/1MC_{\text{tax}} = 0.20 \times \$0.0205 + 0.20 \times 800 \times \$10.00/\text{1M}

Ctax=$0.0041+$0.0016=$0.0057 per query\boxed{C_{\text{tax}} = \$0.0041 + \$0.0016 = \$0.0057 \text{ per query}}

This $0.0057 per-query penalty can be expressed against two different denominators, and both appear in this document — so both are defined here explicitly:

CtaxCgen=$0.0057$0.0205=27.8%(overhead on the generation pass)\frac{C_{\text{tax}}}{C_{\text{gen}}} = \frac{\$0.0057}{\$0.0205} = 27.8\% \quad \text{(overhead on the \textit{generation pass})}

CtaxCtotal=$0.0057$0.0263=21.7%22%(share of total per-query cost)\frac{C_{\text{tax}}}{C_{\text{total}}} = \frac{\$0.0057}{\$0.0263} = 21.7\% \approx 22\% \quad \text{(share of \textit{total per-query cost})}

Where Ctotal=$0.0263C_{\text{total}} = \$0.0263 includes generation, local NLI verification, and the tax itself (§10.1). Both figures describe the same $0.0057; the "22%" used elsewhere in this document always refers to the second.

Important

The Hallucination Tax is not a bug. It is a mechanical consequence of two architectural constraints: (1) proprietary APIs do not expose generation-state control, making partial rollback impossible (Structural), and (2) the post-hoc verification pattern evaluates the response after all tokens are generated and billed (Inferred). RLHF fine-tuning reduces the hallucination rate rr but cannot reduce the cost per event — that cost is fixed by the API boundary, not by model quality. Even if fine-tuning halves the rate from 20% to 10%, each remaining event still forces a complete regeneration cycle.

4.4 Scaling Economics

The Hallucination Tax scales linearly with query volume. Let QQ denote daily query volume:

Cdaily=QCtaxC_{\text{daily}} = Q \cdot C_{\text{tax}}

The two tiers below are defined explicitly, since the multiplier differs by model:

  • Mid-Tier = GPT-4o generation + local cross-encoder NLI → Ctax=$0.0057C_{\text{tax}} = \$0.0057/query
  • High-Tier = Claude Opus 4.8 generation → Ctax=$0.0130C_{\text{tax}} = \$0.0130/query (same formula, Opus pricing: Cgen=$0.045C_{\text{gen}} = \$0.045)
Daily Query Volume Monthly Hallucination Tax (Mid-Tier) Monthly Hallucination Tax (High-Tier)
10,000 queries/day $1,710 $3,900
50,000 queries/day $8,550 $19,500
100,000 queries/day $17,100 $39,000
1,000,000 queries/day $171,000 $390,000

(30-day month. Modeled — query volumes are illustrative scale points, not the vendor's traffic, which I have no visibility into. The defensible figure is the per-query tax; the monthly totals simply show its linearity.)

4.5 Cross-Scenario Comparison

Scenario Generation Model Verification Method Total Per-Query Cost Compliance Overhead†
A GPT-4o (mid-tier) Cross-encoder NLI (local) $0.0263 $0.0057 (22%)
B GPT-4o (mid-tier) LLM-as-Judge $0.0427 $0.0222 (52%)
C Claude Opus 4.8 (high-tier) LLM-as-Judge $0.0920 $0.0470 (51%)

Compliance Overhead = verification call cost + Hallucination Tax. This is a broader quantity than the Hallucination Tax defined in §4.3, and the two must not be conflated. Breakdown: Scenario A = $0.0000 judge + $0.0057 tax. Scenario B = $0.0165 judge + $0.0057 tax. Scenario C = $0.0340 judge + $0.0130 tax. In Scenario A the local NLI cost is negligible (~$0.0001), so overhead ≈ tax; in B and C it dominates.

Warning

In Scenario B (LLM-as-Judge verification), the verification call alone costs $0.0165/query — roughly 80% of the primary generation cost — repeated on every single query. Together with the Hallucination Tax this produces a cost structure where over half the generation-path spend is compliance overhead with no output value.


5. Compliance Gap Analysis

This section names the specific compliance gaps the analysis identifies. These are the diagnostic findings that would appear in a formal compliance assessment — derived from the architecture reconstructed in §3 and §4.1, not from an executed audit of the vendor's systems.

# Gap Severity Framework Reference
G-1 No inference-time control for hallucination interception — the verification gate fires only after the complete response is generated; there is no active compliance control on the generation path during inference Critical EU AI Act Art. 15 (Accuracy); NIST GOVERN 1.1
G-2 Contaminated generation state persists unchecked — hallucinated token sequences remain in the KV cache and condition all downstream tokens; no mechanism exists to detect or eliminate contamination during generation Critical EU AI Act Art. 14 (Human Oversight); ISO 42001 A.6.2.6
G-3 No audit trail for verification decisions — the post-hoc gate produces a binary pass/fail signal with no structured log of what was verified, what was rejected, or why regeneration was triggered Critical ISO 42001 A.6.2.6; TPRM Requirement
G-4 Verification model has a structural reliability ceiling — 512-token cross-encoder input limit causes truncation or fragmentation of Italian legal clauses, degrading verification accuracy on the queries that matter most High EU AI Act Art. 9 (Risk Management)
G-5 Hallucination guarantee not validated under adversarial conditions — marketed hallucination rate not tested against structurally predictable failure classes (complex clauses, nested cross-references, multi-statute queries) High EU AI Act Art. 9; NIST MEASURE 2.1
G-6 No distinction between verified and unverified output in the user interface — the user receives all output with identical confidence formatting regardless of verification status High EU AI Act Art. 14; SRA Principle 2
G-7 No continuous monitoring of hallucination rate or regeneration frequency — the Hallucination Tax is invisible in aggregate dashboards; no alerting threshold for verification failure rate Medium NIST MEASURE 2.1; ISO 42001 A.6.2.6

Note

Gaps G-1 through G-3 are systemic — they are inherent to the architectural pattern (post-hoc verification on proprietary APIs), not to the implementation quality of any particular team. A competent engineering team operating within the constraints of managed API endpoints and standard orchestration frameworks would produce this exact gap profile. That generality is precisely why the profile can be derived from public architectural facts rather than from privileged access: it is a property of the pattern, and it applies to the category, not to one vendor.


6. The Compliance Control — Architecture of the Remediation

Important

Everything from here forward is Specified. This is a target-state reference architecture. It has not been implemented. Component figures (rollback latency, auditor throughput, GPU cost) are Modeled from published benchmarks and pricing for the named components, not measured from an assembled system.

6.1 The Architectural Principle

The remediation does not involve replacing the vector database, modifying the retrieval pipeline, or retraining the model. It refactors the generation path from a post-generation gating pattern to an in-flight, inference-time compliance control that audits each sentence as it is generated.

This requires one prerequisite change: replacing proprietary API endpoints (which do not expose generation-state control) with a self-hosted open model (Llama-3-70B) deployed on dedicated GPU instances fronted by a vLLM/SGLang inference engine. This transition restores architectural access to the KV cache — the generation state — enabling selective truncation during inference.

Warning

This prerequisite is the single largest cost in the proposal, and it is not primarily a technical one. Moving from a managed API to self-hosted GPU inference transfers an entire operational burden onto the platform team: capacity planning, GPU availability, model upgrades, inference-engine version churn, on-call for a stateful serving tier, and an open-weights model whose Italian-legal-language quality must be independently qualified against the frontier model it replaces (§15.2). The economic model in §10 counts the GPU-hour savings. It does not count this. Any honest reading of the proposal must weigh both.

6.2 The In-Flight Verification Loop

The compliance control operates as a real-time state machine on the generation path. As the model streams tokens, an independent, lightweight 8B sentence-level NLI auditor monitors the generation buffer:

sequenceDiagram
    autonumber
    actor User as User Query
    participant Router as Query Router
    participant Retriever as Retrieval Tier
    participant LLM as Llama-3-70B<br/>(KV Cache Exposed)
    participant Aud as 8B NLI Compliance<br/>Auditor
    participant Log as Audit Trail<br/>(Structured JSON)

    User->>Router: Submit legal query
    Router->>Retriever: Route to retrieval pipeline
    Retriever->>LLM: Deliver evidentiary chunks + query

    loop Sentence-Level Generation & Audit
        LLM->>Aud: Stream sentence buffer
        Note over Aud: Sentence compared against<br/>retrieved evidentiary sub-spans<br/>(ColBERT coordinate extraction)

        alt Entailment (Claim verified)
            Aud->>LLM: Lock KV cache state & continue
            Aud->>Log: Log: ENTAILMENT — sentence verified,<br/>source sub-span hash, confidence score
        else Contradiction (Hallucination detected)
            Aud->>LLM: TRUNCATE KV cache to last<br/>verified token index
            Aud->>LLM: Inject corrective context prefix
            Aud->>Log: Log: CONTRADICTION — sentence rejected,<br/>rollback depth, corrective prefix
            Note over LLM: Generation resumes from<br/>clean state — contaminated<br/>tokens eliminated
        else Neutral (No factual claim)
            Aud->>LLM: Allow passage without lock
            Aud->>Log: Log: NEUTRAL — conversational<br/>sentence, no verification required
        end
    end

    LLM->>User: Deliver verified output
    Log->>Log: Finalise query audit record<br/>with SHA-256 hash

6.3 The KV Cache Truncation Mechanism

When the 8B auditor detects a contradiction — an ungrounded claim — it halts generation immediately and performs a selective rollback by truncating the KV cache back to the index of the last verified sentence. This is the core mechanism that breaks the Verification Trap:

Step 1: The model generates sentence S1S_1. The auditor compares S1S_1 against the retrieved evidentiary sub-spans (extracted via ColBERT token-level coordinates, not the full 500-word chunk). The auditor returns ENTAILMENT. The KV cache coordinates for S1S_1 are locked in the inference engine using a RadixAttention tree fork.

Step 2: The model generates sentence S2S_2, which contains a fabricated citation. The auditor compares S2S_2 against the evidentiary sub-spans and returns CONTRADICTION.

Step 3: The inference engine truncates the active KV cache, discarding all tokens representing S2S_2. The memory state of S2S_2 is deleted — it cannot contaminate downstream generation. This is the operation that proprietary APIs do not expose.

Step 4: The engine injects a corrective context prefix (e.g., "Correction: The retrieved sources do not support the existence of the cited holding. Resuming with verified evidentiary basis...") and resumes generation from the clean state.

The modeled latency cost of this operation versus full regeneration:

trollback+tcorrective120msvs.tfull_regeneration8,00012,000ms\boxed{t_{\text{rollback}} + t_{\text{corrective}} \approx 120\text{ms} \quad \text{vs.} \quad t_{\text{full\_regeneration}} \approx 8{,}000\text{–}12{,}000\text{ms}}

The waste per hallucination event — measured in tokens generated and discarded — drops by 95%:

Tokens wastedrollback40vs.Tokens wastedfull_regen800\text{Tokens wasted}_{\text{rollback}} \approx 40 \quad \text{vs.} \quad \text{Tokens wasted}_{\text{full\_regen}} \approx 800

Tip

What the 95% actually rests on. The token counts are the honest part of this claim and they are close to arithmetic: a sentence-boundary catch discards one sentence (~40 tokens); a post-hoc catch discards a whole response (~800 tokens). That ratio follows from where the boundary sits, not from any performance tuning — it is the clearest expression of the structural argument. The 120ms rollback figure is softer: it is modeled from RadixAttention tree-navigation costs and assumes the auditor keeps pace with the generation stream, which is exactly the assumption §15.1 flags as most likely to break under concurrency.

6.4 The Three-State Compliance Machine

The auditor operates on a three-state transition model designed to handle edge cases without entering infinite verification loops:

sequenceDiagram
    participant M as Model
    participant S as System
    participant A as Auditor
    participant UI as User Interface

    M->>S: Produces sentence buffer
    S->>A: Request verification

    alt Verdict: ENTAILMENT
        A-->>S: ENTAILMENT
        Note over S: Lock KV cache state
        S->>UI: Stream verified text
        S->>M: Continue generation

    else Verdict: CONTRADICTION
        A-->>S: CONTRADICTION
        Note over S: Check rollback count for sentence

        alt Count ≤ 3
            Note over S: Truncate KV Cache
            S->>M: Inject corrective context prefix
            M-->>S: Resume generation from clean state

        else Count > 3
            S->>A: Escalate to frontier model re-audit
            S->>UI: Mark sentence as "Unverified"
        end

    else Verdict: NEUTRAL
        A-->>S: NEUTRAL
        Note over S: Allow passage (no factual claim)
        S->>M: Continue generation
    end
State Trigger Action Compliance Implication
Entailment Claim verified against retrieved evidentiary sub-spans Lock KV cache state, stream to user, continue Output is verified and auditable — audit trail records the source sub-span and confidence score
Contradiction Claim violates retrieved evidence Truncate KV cache, inject corrective prefix, resume Contaminated state eliminated before downstream conditioning; event logged for incident forensics
Neutral Sentence contains no factual assertion (e.g., "Based on the contracts analysed...") Allow passage without locking KV cache Conversational tokens pass through; no verification claim is made in the audit trail
Escalation Contradiction count exceeds 3 for a single sentence Escalate to frontier model re-audit OR append "Unverified" label in UI Prevents infinite loops while maintaining transparency; user is informed when verification is inconclusive

Note

The rollback ceiling of 3 is a design parameter, not a tuned value — it bounds worst-case latency at four generation attempts per sentence. The correct value is an empirical question that only a running system can answer, and it trades user-visible "Unverified" labels against tail latency. It is called out here rather than presented as settled.

6.5 The ColBERT Sub-Span Extraction

A critical implementation detail: the 8B auditor does not receive the full 500-word retrieved chunk as its verification premise. Full chunks routinely exceed the NLI model's effective context window and introduce noise that degrades verification precision.

Instead, ColBERT's token-level MaxSim scores are used to extract the exact ~50-word sub-span from the retrieved chunk that supports (or contradicts) the generated claim:

SubSpan(Si,Cj)=argmaxspanCjkspanmaxlSicos(ck,sl)\text{SubSpan}(S_i, C_j) = \arg\max_{\text{span} \subseteq C_j} \sum_{k \in \text{span}} \max_{l \in S_i} \cos(\mathbf{c}_k, \mathbf{s}_l)

This sub-span extraction serves two purposes:

  1. Verification precision: The auditor compares a ~40-token sentence against a ~50-token evidentiary sub-span — well within the effective range of the 8B NLI model.
  2. Audit trail granularity: The log records exactly which sub-span of which source document was used to verify which sentence — providing the retrieval-to-generation traceability that TPRM evaluators require.

6.6 Architecture Comparison: Current vs. Target State

graph LR
    subgraph PRE ["Current State (Uncontrolled)"]
        direction LR
        Q1([User Query]) --> R1[Retrieve<br/>Chunks]
        R1 --> G1["Generate Full<br/>Response<br/>(Proprietary API)"]
        G1 --> PH["Post-Hoc<br/>Verification Gate"]
        PH -->|Pass| U1([User])
        PH -->|Fail| G1
    end

    subgraph POST ["Target State (Specified)"]
        direction LR
        Q2([User Query]) --> R2[Retrieve<br/>Chunks]
        R2 --> G2["Generate with<br/>In-Flight Audit<br/>(Self-Hosted)"]
        G2 --> CG["Per-Sentence<br/>Compliance Gate"]
        CG -->|Verified| U2([User])
        CG -->|Contradiction| TB["Truncate &<br/>Resume"]
        TB --> G2
    end

7. Compliance Outcome Table

Caution

The right-hand column is a design target, not a measurement. No row below has been verified against a running system. The "How It Would Be Verified" column is the test that would establish each claim — it is the work still to be done, and stating it that way is the point of the column.

Compliance Outcome Current State Target State (Specified) How It Would Be Verified Basis
Contaminated output reaching end users Uncontrolled — hallucinated tokens pass gate or trigger full regen 0% of detected contamination reaches output Adversarial battery + shadow parity comparison Structural, scoped — truncation provably removes detected contamination from generation state; it cannot remove what the auditor fails to detect (§15.1)
Generation-tier audit trail None — binary pass/fail signal only Full — structured JSON log per query with per-sentence decisions, source sub-span references, tamper-evident hashing Schema conformance test against §8 spec Structural — the log is emitted by the control by design
Inference-time compliance control active No — verification fires only post-generation Yes — in-flight sentence-level auditor operates during generation Architecture validation + shadow parity testing Specified — not built
Verification model reliability ceiling 512-token cross-encoder limit causes truncation/fragmentation Raised — 8B auditor with ColBERT sub-span extraction operates within effective context range Sub-span extraction benchmarked against full-chunk baseline Modeled — the token-budget argument is sound; the precision gain needs measuring
User distinction between verified/unverified output None — all output presented identically Active — sentences exceeding rollback threshold labelled "Unverified" in UI UI audit + escalation logic validation Structural — follows from the escalation path in §6.4
Continuous monitoring capability None — hallucination rate invisible in aggregate dashboards Active — per-query rollback rate, escalation frequency, verification confidence tracked Dashboard build + alerting threshold validation Specified — the log carries the required fields; the dashboard does not exist
EU AI Act Article 15 conformity No accuracy control on generation path Materially strengthened — deterministic inference-time control with auditable outcomes Regulatory mapping + structured test protocol; formal assessment by qualified assessor Specified — an architecture cannot self-certify conformity
TPRM retrieval accuracy section Fails — no measurable generation accuracy control Answerable — auditable control + audit trail + monitoring Compliance outcome documentation Specified

Secondary: Projected Performance & Economic Outcomes

All rows Modeled. See §10 for the cost derivation and §15 for what would invalidate it.

Performance Outcome Current State Target State (Modeled)
User-facing latency (hallucination event) 12–25 seconds (full regeneration) 1.8–2.8 seconds (120ms rollback, seamless to user)
Waste per hallucination event ~800 tokens discarded + full regen ~40 tokens discarded (95% reduction)
Per-query Hallucination Tax (mid-tier) $0.0057 ~$0.00002
Total per-query cost (mid-tier) $0.0263 $0.00205 (92.2% reduction)
Total per-query cost (high-tier + LLM-as-Judge) $0.0920 $0.00205 (97.8% reduction)
Monthly compute waste (50K queries/day, mid-tier) $8,550 ~$30

8. Audit Trail & Measurability Specification

The in-flight compliance control emits a structured JSON audit record for every query processed. This log is designed to satisfy four compliance functions: decision traceability, anomaly detection, incident forensics, and continuous compliance monitoring.

8.1 Specimen Audit Log Entry

Note

The following is a specimen record conforming to the specified schema — an illustration of the log format with representative values. It is not an extract from a running system, and the Italian legal text and hashes within it are illustrative.

{
  "query_id": "q-20251215-004712",
  "timestamp": "2025-12-15T14:47:12.331Z",
  "session_id": "sess-a3f8c921",
  "query_hash": "sha256:e4b2c1...",
  "generation_model": "llama-3-70b",
  "auditor_model": "nli-8b-legal-v2",
  "retrieval_context": {
    "chunks_retrieved": 4,
    "total_tokens": 3200,
    "source_ids": ["it-cc-art42", "it-cc-art2359", "eu-gdpr-art82", "it-dlgs-231-art6"]
  },
  "sentence_audit_trail": [
    {
      "sentence_index": 0,
      "text": "L'articolo 42 del Codice Civile stabilisce i requisiti per la costituzione...",
      "verdict": "ENTAILMENT",
      "confidence": 0.94,
      "source_subspan": {
        "chunk_id": "it-cc-art42",
        "token_range": [12, 67],
        "subspan_text": "requisiti per la costituzione delle persone giuridiche..."
      },
      "kv_cache_action": "LOCKED",
      "latency_ms": 18.4
    },
    {
      "sentence_index": 1,
      "text": "La responsabilità solidale si applica ai sensi del decreto...",
      "verdict": "CONTRADICTION",
      "confidence": 0.89,
      "source_subspan": {
        "chunk_id": "it-cc-art42",
        "token_range": [68, 112],
        "subspan_text": "nessuna disposizione relativa alla responsabilità solidale..."
      },
      "kv_cache_action": "TRUNCATED_TO_INDEX_0",
      "corrective_prefix_injected": true,
      "rollback_count": 1,
      "latency_ms": 22.1
    },
    {
      "sentence_index": 2,
      "text": "Pertanto, i requisiti formali includono l'atto costitutivo...",
      "verdict": "ENTAILMENT",
      "confidence": 0.91,
      "source_subspan": {
        "chunk_id": "it-cc-art42",
        "token_range": [14, 58],
        "subspan_text": "l'atto costitutivo e lo statuto devono contenere..."
      },
      "kv_cache_action": "LOCKED",
      "latency_ms": 16.7
    }
  ],
  "summary": {
    "total_sentences": 8,
    "entailment_count": 6,
    "contradiction_count": 1,
    "neutral_count": 1,
    "rollback_events": 1,
    "escalation_events": 0,
    "total_generation_latency_ms": 2340,
    "total_audit_latency_ms": 147
  },
  "human_intervention_required": false,
  "audit_hash": "sha256:7f3a91c2d8e4b..."
}

8.2 What the Log Is Designed to Satisfy

Compliance Function How the Log Addresses It
Decision traceability Every sentence maps to a specific verdict, a specific source sub-span, and a specific KV cache action — the complete chain from retrieval to generation to verification is logged
Anomaly detection Rollback rate, escalation frequency, and confidence distribution are tracked per query — spikes indicate retrieval degradation or model drift
Incident forensics If a user reports a fabricated citation, the log provides the exact sentence, the exact source sub-span compared, the auditor verdict, and whether the sentence was rolled back or passed through
Continuous compliance The log feeds a monitoring dashboard with alerting thresholds for rollback rate (>30% triggers investigation), escalation rate (>5% triggers model review), and confidence distribution drift

Tip

The audit_hash field is a SHA-256 hash of the complete audit record, making the entry tamper-evident — the decision trace can be proven unmodified after the fact. Note the scope precisely: a self-computed hash proves integrity against later alteration; it is not a trusted timestamp and does not by itself prove when the record was created. Anchoring to an external notary or append-only log is the next step if that property is required.


9. Regulatory Applicability Matrix

This matrix maps the specified compliance control to the regulatory provisions it is designed to address. It pairs the gap analysis from §5 with the proposed remediation.

Caution

"Addressed by Design" is not "Compliant." Every status below describes what the architecture is designed to satisfy once built and validated. None represents an executed conformity assessment, a certification, or a regulator's determination. Conformity is assessed by qualified assessors against a running system — never claimed by its architect.

Regulatory Framework Provision Gap Identified (§5) Control Specified Status
EU AI Act Art. 15 — Accuracy G-1: No inference-time accuracy control In-flight sentence-level NLI auditor with KV cache truncation Addressed by design — pending build and validation
EU AI Act Art. 14 — Human oversight G-2: No mechanism to intercept contamination; G-6: No verified/unverified distinction Three-state machine with "Unverified" UI labelling on escalation Addressed by design — pending build and validation
EU AI Act Art. 9 — Risk management G-5: Hallucination guarantee not adversarially validated Structured test protocol for generation-tier failure classes Protocol defined — test suite specified, not executed
Italian Codice Deontologico Art. 9 & 12 — Diligence G-1, G-6: Fabricated holdings reach user indistinguishably Verified output labelled; unverified output flagged Addressed by design — pending build and validation
SRA Principle 2 — Public trust G-1: Hallucinated text presented as authoritative Contaminated state eliminated during inference Addressed by design, scoped to detected contamination
NIST AI RMF GOVERN 1.1 G-3: No audit trail; G-7: No monitoring Structured JSON audit log + monitoring dashboard Schema specified — dashboard not built
ISO 42001 A.6.2.6 G-3, G-7: No quality objectives defined or monitored Per-query verification metrics with alerting thresholds Objectives defined — not yet monitored

10. Unit Economics / OPEX Impact

Note

This section is Modeled. The pre-remediation column derives from published API pricing (§4.2); the post-remediation column derives from published GPU instance pricing amortised over a declared assumption of sustained high utilisation. That utilisation assumption is the model's softest input and is examined in §15.3.

10.1 Per-Query Cost Comparison

The specified architecture replaces proprietary API billing with self-hosted GPU inference, eliminating both the per-token billing model and the Hallucination Tax:

Cost Component Current State (GPT-4o + NLI) Target State (Llama-3-70B + 8B Auditor)
Primary generation $0.0205/query (per-token API) $0.00175/query (GPU-hour amortised)
Verification ~$0.0001 (local NLI) $0.000135 (8B auditor on L4 GPU)
Hallucination Tax (regen + waste) $0.0057/query ~$0.00002 (40-token rollback)
Heavy model escalation (<5%) $0.00015
Total per query $0.0263 $0.00205

Modeled per-query cost reduction=$0.0263$0.00205$0.0263=92.2%\boxed{\text{Modeled per-query cost reduction} = \frac{\$0.0263 - \$0.00205}{\$0.0263} = 92.2\%}

Warning

Two costs this table does not carry. First, GPU amortisation assumes sustained utilisation; at low or spiky traffic the fixed cost per query rises sharply and can exceed the API baseline outright (§15.3). Second, self-hosting adds engineering and on-call cost that no per-query figure captures (§6.1). The 92.2% is a compute-cost reduction, not a total-cost-of-ownership reduction, and it should not be quoted as one.

10.2 Hallucination Tax Elimination

The Hallucination Tax — the specific per-query OPEX penalty created by post-hoc verification on proprietary APIs — is modeled to fall by over 99%:

Ctax,beforeCtax,after=$0.0057$0.00002=285×\frac{C_{\text{tax,before}}}{C_{\text{tax,after}}} = \frac{\$0.0057}{\$0.00002} = 285\times

The mechanism is the honest core of this claim: each hallucination event wastes ~40 tokens (caught at sentence boundary, cache truncated) instead of ~800 tokens (full response generated, verified, discarded, regenerated). The model still hallucinates at the same underlying rate — the cost per event is what changes.

10.3 Scaled Economics

Scale Monthly Cost (Current) Monthly Cost (Target) Monthly Delta
10K queries/day $7,890 $615 $7,275
50K queries/day $39,450 $3,075 $36,375
100K queries/day $78,900 $6,150 $72,750

(30-day month, Modeled. At the 10K/day row especially, the GPU utilisation assumption in §15.3 deserves scrutiny before the delta is treated as real.)


11. Deployment Specification: Zero-Downtime Parallel Validation

Important

This is a migration protocol, not a migration report. No phase below has been executed. There is no shadow cluster, no 72-hour soak, and no telemetry. The figures in Phase 2 are acceptance gates — thresholds a team should require before cutting over — stated as targets, never as results.

The migration would be executed as a parallel validation deployment (shadow parity) to eliminate engineering risk:

Phase 1: Dual-Path Shadow Auditing

The existing production application continues to run on managed API endpoints. Production traffic is asynchronously mirrored to a staging cluster running the self-hosted Llama-3-70B + in-flight auditor. No production user sees staging output.

graph TD
    User([Production Traffic]) --> Prod["Production Pipeline<br/>(Proprietary API +<br/>Post-Hoc Gate)"]
    User -.->|"Async mirror<br/>(100% traffic)"| Shadow["Shadow Pipeline<br/>(Self-Hosted +<br/>In-Flight Auditor)"]

    Prod --> ProdOut([Production Output<br/>to User])
    Shadow --> Telemetry["Telemetry Comparison<br/>━━━━━━━━━━━━━━━━<br/>• Accuracy delta<br/>• Latency delta<br/>• Rollback rate<br/>• Contamination caught"]

    style Prod fill:#000,stroke:#f57c00,stroke-width:2px
    style Shadow fill:#000,stroke:#1a73e8,stroke-width:2px
    style Telemetry fill:#000,stroke:#16a34a,stroke-width:2px

Phase 2: Telemetry Validation — Acceptance Gates

A minimum 72-hour soak, with cutover gated on the criteria below. Each gate has a threshold and a defined failure action:

# Gate Threshold to Proceed If the Gate Fails
1 Contamination interception rate — share of production-leaked hallucinations the shadow pipeline catches and corrects ≥ 90% against a human-adjudicated sample The 8B auditor's recall is the bottleneck; qualify a larger auditor or improve sub-span extraction before proceeding
2 Auditor false-positive rate — verified content incorrectly flagged as contradiction < 1% of sentences Rollback loops will degrade fluency and inflate latency; retune thresholds or the NLI model
3 p99 latency, hallucination-triggered path < 3s (vs. modeled 12–25s baseline) The 120ms rollback assumption (§6.3) does not hold under real concurrency; re-model §7 secondary table
4 Generation quality parity — Llama-3-70B vs. incumbent frontier model on Italian legal output, blind-rated by qualified reviewers No material regression This is the gate most likely to fail and the least discussed (§15.2). A quality regression can invalidate the whole migration regardless of cost savings.
5 Escalation rate — sentences exceeding the rollback ceiling < 5% Either the retrieval tier or the rollback ceiling (§6.4) needs work
6 GPU utilisation under real traffic shape Sufficient to hold the §10.1 amortised cost The economic case weakens or inverts; re-run §10 with measured utilisation

Note

Gate 4 exists because it is the failure mode a cost-driven proposal is most tempted to omit. Every other gate measures the new control; Gate 4 measures what was given up to enable it. An architecture that halves compute cost while degrading Italian legal reasoning quality is a worse system, and no audit trail compensates for that.

Phase 3: Controlled Rollout with Rollback Capability

A feature flag routes a progressively increasing share of traffic to the new pipeline. The managed-API path is retained — not decommissioned — through a defined bake period, so rollback remains a flag flip rather than a re-migration. Target production downtime: zero.


12. Technical Deep Dive — Appendix

12.1 Why Proprietary APIs Create the Structural Lock-In

This is the load-bearing argument of the document. It is Structural — evaluable from the semantics of the API boundary alone, with no benchmark, no deployment, and no vendor-specific information required.

The Hallucination Tax is not a feature of any specific platform's implementation. It is a mechanical consequence of the interaction between three architectural properties:

  1. Autoregressive generation: Every token is conditioned on the full preceding sequence. Contamination at token tkt_k propagates to all tokens tk+1,tk+2,,tnt_{k+1}, t_{k+2}, \dots, t_n.

  2. Opaque inference endpoints: Proprietary APIs (OpenAI, Anthropic, Google) expose a request/response interface. The application sends a prompt and receives a completed response. The KV cache — the internal state that stores attention computations for all generated tokens — is not accessible to the application. The application cannot inspect, modify, or selectively truncate this state during generation.

  3. Per-token billing on all generated tokens: Managed APIs bill for every output token generated, regardless of whether the application uses the response. When a post-hoc gate rejects a response and triggers regeneration, the wasted output tokens on the discarded response are billed at full rate.

These three properties interact to create a structurally irreducible cost penalty: the application cannot prevent contamination (property 1), cannot remediate contamination mid-generation (property 2), and pays full cost for contaminated output (property 3).

This is why the remediation requires self-hosting rather than better engineering on the existing stack. No amount of implementation quality removes property 2 from a managed endpoint. That is the claim this specification actually proves, and it is independent of every modeled number elsewhere in the document.

Note

Scope of the claim. Property 2 describes managed endpoints as they expose functionality today. It is a statement about an interface contract, not a law of nature — a provider could ship a partial-rollback or constrained-resume primitive tomorrow and collapse this gap. The architectural point would survive: the control must live where the generation state lives.

12.2 RadixAttention Tree Fork for KV Cache Management

The selective KV cache truncation mechanism uses a RadixAttention tree structure (as implemented in SGLang) to manage multiple generation branches. Each verified sentence creates a branch point in the radix tree:

RadixTree:Token SequenceKV Cache Coordinates\text{RadixTree}: \text{Token Sequence} \to \text{KV Cache Coordinates}

When a rollback occurs, the engine navigates the radix tree back to the last locked branch point and resumes generation from that state. Tokens on the discarded branch are garbage-collected. This is functionally equivalent to a "selective forget" operation — the model's memory of the hallucinated content is mechanically deleted.

12.3 Why RLHF Cannot Substitute for Inference-Time Control

RLHF fine-tuning and inference-time auditing operate on orthogonal axes:

Dimension RLHF Fine-Tuning In-Flight Auditing
What it addresses Hallucination rate (rr) Hallucination cost per event (CeventC_{\text{event}})
When it operates Training time Inference time
Total cost formula Ctotal=QrCeventC_{\text{total}} = Q \cdot r \cdot C_{\text{event}} Ctotal=QrCeventC_{\text{total}} = Q \cdot r \cdot C_{\text{event}}
What it reduces rr (from 20% to perhaps 10%) CeventC_{\text{event}} (from $0.02 to $0.0001)
Structural limit Cannot reach r=0r = 0; diminishing returns Reduces CeventC_{\text{event}} substantially regardless of rr

The interventions are complementary, not competing. Fine-tuning reduces how often hallucinations occur. In-flight auditing reduces how much each hallucination costs. The optimal architecture uses both.


13. Lessons Learned

1. Post-hoc verification is a compliance control that fires after the damage has propagated.

The standard "generate-then-validate" pattern creates a compliance-free zone between the first hallucinated token and the moment the verification gate evaluates the complete response. In that interval, contaminated attention states accumulate in the KV cache and condition every subsequent token. For regulated deployments where output accuracy carries professional liability, the verification boundary must operate during generation — not after it. A compliance control that can only detect contamination after it has propagated through the entire output is structurally insufficient for high-stakes domains.

2. Proprietary API boundaries create an architectural ceiling on compliance control granularity.

Managed inference endpoints optimise for developer velocity — they abstract away the inference engine's internal state behind a clean request/response interface. This abstraction is the correct trade-off for general-purpose applications. In regulated domains requiring inference-time output control, the same abstraction becomes a structural constraint: the application cannot implement compliance controls that require access to the generation state. Compliance audits that do not assess the inference-time control surface available on the chosen model hosting architecture are incomplete by design.

3. The Hallucination Tax is a mechanical cost, not a model quality problem.

The per-event cost of hallucination — full regeneration, wasted tokens, re-processed input — is determined by the API billing model and the verification architecture, not by the model's hallucination rate. RLHF fine-tuning reduces how often the tax is triggered but cannot reduce the tax amount. For regulated deployments at scale, the cost per hallucination event — not just the hallucination rate — must be an explicit term in the unit economics model.

4. Generation-tier compliance controls require retrieval-tier coordination.

The in-flight auditor's effectiveness depends on the quality of the evidentiary sub-spans it receives from the retrieval tier. Passing the full 500-word retrieved chunk to an NLI model degrades verification precision. ColBERT token-level coordinate extraction — selecting the specific ~50-word sub-span that supports or contradicts the generated claim — is not an optimisation; it is a prerequisite for reliable sentence-level verification. Compliance controls on the generation path cannot be designed in isolation from the retrieval path.

5. This failure class is structurally invisible to standard engineering practice.

The post-hoc verification pattern is the default architecture in every major RAG orchestration framework. A competent team following standard practices — retrieval, generation, validation — would produce this exact gap profile. The Hallucination Tax is invisible in aggregate cost dashboards (it looks like normal API spend) and invisible in aggregate quality dashboards (the gate catches most contamination — eventually). Surfacing this failure class requires specialised compliance testing that measures the cost and latency of the verification loop itself, not just the quality of the output.

6. An architectural argument and a benchmark result are different instruments, and conflating them destroys both.

The strongest claim in this document — that partial remediation is impossible behind a managed API boundary — has no measurement behind it and needs none; it follows from the interface contract. The most impressive-looking claims (92.2% cost reduction, 285× tax elimination, 120ms rollback) are the weakest, because each rests on declared assumptions about tokens, utilisation, and concurrency. Presenting both under one undifferentiated "Results" heading would have made this document more persuasive and less true. In a domain where buyers conduct technical due diligence before purchase, a claim that cannot survive the follow-up question is worth less than no claim at all.


14. What This Specification Does Not Establish

Stated plainly, so no reader has to infer it from what is missing:

  • It has not been built. No GPU cluster was provisioned, no auditor was trained or qualified, no shadow migration was run, and no query has ever passed through this control.
  • The 0% contaminated-output target is scoped to detected contamination. Truncation provably removes what the auditor catches. It does nothing about what the auditor misses, and an 8B NLI model on Italian legal text will miss things. The honest framing is a large reduction in contamination reaching users, with a residual rate set by auditor recall — not zero contamination in the absolute.
  • The diagnosis is a reconstruction from public sources. I had no access to the analysed platform. If its actual architecture differs from the §4.1 reconstruction, the diagnosis does not apply to it — though the failure class remains real for systems built on framework defaults.
  • The Stanford 17–33% rate was measured on other systems. It is Confirmed research, applied here by analogy. It is not a measurement of this platform.
  • No regulatory conformity is claimed or certified. §9 states design intent against provisions. Conformity assessment is performed by qualified assessors against running systems.
  • The economics are compute-only. Engineering time, on-call burden, GPU procurement risk, and model qualification effort are real costs this model does not carry (§6.1, §10.1).

15. Residual Risks & What Would Falsify This

The architectural argument in §12.1 holds. The engineering around it has genuine soft spots, and they are not where the cost model draws attention.

15.1 Auditor Recall Is the Real Ceiling

The entire control is downstream of one question: does the 8B NLI auditor actually detect the contradiction? Every guarantee in §7 is scoped to detected contamination, which makes auditor recall the binding constraint on the whole design.

  • Legal entailment is harder than general-domain NLI. A subtly misattributed holding may be textually consistent with the retrieved sub-span while being legally wrong. NLI models are trained to detect contradiction, not misattribution.
  • The 8B auditor is smaller than the generator it polices. There is an inherent asymmetry in asking a smaller model to catch a larger model's sophisticated errors.
  • False positives have their own cost. An over-eager auditor truncates valid reasoning, triggers rollback loops, degrades fluency, and inflates latency — Gate 2 in §11 exists for exactly this.

15.2 The Model Substitution Is Not Cost-Neutral

Replacing a frontier API model with Llama-3-70B is presented in §10 as a cost line. It is also a capability change on Italian-language legal reasoning, and that direction is not guaranteed favourable. If output quality regresses materially, the control is protecting a worse product — and no audit trail compensates for that. This is Gate 4 in §11 and it is the gate most likely to fail.

15.3 GPU Amortisation Assumes Utilisation That May Not Exist

The $0.00175/query generation cost holds only under sustained high utilisation. Legal research traffic is bursty and business-hours-shaped. Under a realistic diurnal pattern with idle overnight capacity, effective per-query cost can rise several-fold, and at low volume the self-hosted path can cost more than the API baseline it replaces. The break-even volume is the number a buyer should ask for first — and it is not in this document, because computing it honestly requires the vendor's real traffic shape.

15.4 Sentence-Level Auditing Under Concurrency

The 120ms rollback figure assumes the auditor keeps pace with the generation stream. Under concurrent load, auditor inference contends with generation for GPU resources. The auditing tier may need independent scaling, which changes the cost model in §10 in a direction it does not currently account for.

15.5 What Would Falsify the Core Claims

Observation What It Would Mean
A managed API ships a partial-rollback or constrained-resume primitive §12.1's structural lock-in dissolves; the self-hosting prerequisite — and most of the cost case — evaporates
8B auditor recall on legal contradiction measures poorly The control's central guarantee is scoped near-vacuously; a larger auditor is required and the cost model must be re-run
Auditor false-positive rate exceeds a few percent Rollback loops make the system slower and less fluent than the post-hoc baseline — a net regression
Real traffic utilisation falls below break-even §10 inverts; the managed API is the cheaper architecture and the proposal should be rejected on its own economics
Llama-3-70B underperforms on Italian legal reasoning The migration trades accuracy for cost — the exact trade this document argues against making

16. TL;DR

The diagnosis (reconstructed from public sources). A legal-RAG platform's generation path ships fabricated legal text to practising lawyers on a structurally predictable basis — a compliance control deficiency, not a model quality issue. A post-hoc verification gate evaluates responses after all tokens are generated and billed; when it catches contamination, it forces a full regeneration cycle costing 12–25 seconds and paying for the discarded output. Modeled at published pricing, this is a ~22% penalty on total per-query cost — the "Hallucination Tax" — compounding linearly with query volume.

The structural claim (provable, and the actual point). No amount of model scaling, prompt engineering, or RLHF fine-tuning reduces the cost per hallucination event, because that cost is fixed by the proprietary API boundary: managed endpoints do not expose the generation state required for partial remediation. This follows from the interface contract and needs no benchmark.

The specification (designed, not built). An in-flight sentence-level compliance control with selective KV cache truncation that eliminates contaminated state during inference — modeled to reduce waste per hallucination event by ~95% (~800 tokens → ~40) while producing a structured, tamper-evident audit trail per query. It requires self-hosting, which carries operational and model-quality costs the compute model does not capture (§15).

The generalised principle: in regulated domains, the verification boundary must operate during generation — post-hoc gating is a compliance control that fires after the damage has propagated.

Status: reference architecture. No vendor engagement, no privileged access, no deployment, no telemetry.

The Impact

This specification defines an in-flight generation-tier compliance control that audits each sentence during inference — before downstream tokens are conditioned on ungrounded sequences. Contaminated token sequences are mechanically eliminated from the generation state via selective cache truncation. The control operates on a three-state verification machine (entailment / contradiction / neutral) with deterministic escalation logic, producing a structured audit trail per query that addresses TPRM retrieval accuracy, decision traceability, and continuous monitoring requirements. The architectural argument — that partial rollback is impossible behind a managed API boundary and possible with an exposed KV cache — is provable by construction. The performance and cost figures are analytical models built on published pricing and stated assumptions. The control has not been built or deployed.

Run this measurement against your own system.

Deployment Audit: £500, fixed scope. Credited in full against the next stage.

What this costs