An end-to-end, agentic Legal NLP system that turns raw regulatory text into atomic, traceable requirements and checks any target document against them — with retrieved evidence, calibrated confidence, and human-interpretable justification for every decision.
Figure 1. The ARCCS pipeline. The Regulatory Processing & Extraction Module (RPEM) turns a raw regulation into atomic, traceable requirements; the Compliance Classification Module (CCM) checks a target document against each one, returning a label, a confidence score, and cited evidence.
Regulatory compliance checking — deciding whether a target document satisfies the obligations of a regulation — requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications.
This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. Rather than forcing a verdict on every requirement, ARCCS assigns one of four labels — Compliant, Non-Compliant, Insufficient Information, and Human Required — making abstention and human deferral explicit outcomes.
On an EU public-procurement benchmark of more than 1,200 individual rule checks, ARCCS attains 98.8% accuracy in violation detection, and an LLM-as-a-judge evaluation finds its justifications legally and evidentially consistent in up to 96.67% of the assessed cases. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.
ARCCS takes a regulatory document D and a target document T and returns a structured report stating — for every obligation in D — whether T satisfies it, with what confidence, and on the basis of which textual evidence.
Converts a raw regulation into a compact, non-redundant set of atomic requirement objects.
(f, α, σ, M, src): legal function, regulated actor, scope, mandatory conditions, and a pointer back to the source text.Evaluates each applicable requirement against the target document in three steps.
ARCCS never forces a verdict. With confidence thresholds θlo < θhi, every requirement resolves to one of four outcomes:
Every mandatory condition is explicitly supported by the evidence, nothing contradicts the requirement, and confidence κ ≥ θhi.
The evidence directly contradicts a mandatory condition, or omits one while addressing its topic, with high confidence.
The evidence references the topic but doesn't settle the mandatory conditions — principled abstention rather than a guess.
Ambiguity, contradictory evidence, or confidence below θlo — the case is explicitly routed to a human expert.
A realistic setting with no deterministic gold standard (GDPR vs. real terms-of-use documents, judged by LLM-as-a-judge), and a controlled setting with deterministic ground truth (an EU public-procurement event-log benchmark).
| Agreement | κ | p-value |
|---|---|---|
| Cohen's κ (5.1 vs 5.2) | 0.47 | 0.19 |
| Cohen's κ (5.1 vs 5.2-pro) | 0.47 | 0.19 |
| Cohen's κ (5.2 vs 5.2-pro) | 1.00 | < 0.001 |
| Fleiss' κ (3 raters) | 0.69 | < 0.001 |
ARCCS autonomously flagged the WhatsApp Terms of Service as conflicting with GDPR's accuracy and transparency principles. These very terms were independently found non-compliant by the Irish Data Protection Commission, which fined WhatsApp €225M under Articles 5(1)(a) and 12–14 GDPR. Similarly, in the ChatGPT Terms of Use, ARCCS flags Article 79 (effective judicial remedy) as Non-Compliant: the mandatory-arbitration / San Francisco-forum clauses conflict with EU/EEA data subjects' right to sue in their Member State of residence.
| # | Rule | Dimension |
|---|---|---|
| 1 | Maximum contract-amount threshold | Monetary |
| 2 | Prohibition of duplicate publication of the same call | Duplicate publication |
| 3 | Prohibition of award before publication | Temporal / logical |
| 4 | Prohibition of award before participation | Temporal / logical |
| 5 | Publication and participation must eventually lead to an award | Temporal / logical |
| 6 | Maximum delay of 70 days between publication and award | Temporal / logical |
| 7 | Prohibition of contract start without prior award | Lifecycle |
| 8 | Prohibition of contract end before publication | Lifecycle |
| 9 | Prohibition of contract end right after publication, skipping participation/award | Lifecycle |
| 10 | Any started contract must eventually end | Lifecycle |
| 11 | Any ended contract must previously have started | Lifecycle |
| 12 | Any terminated contract must previously have been awarded | Lifecycle |
| Model | Acc. | Prec. | Rec. | F1 |
|---|---|---|---|---|
| gpt-5.4 | 98.5 | 83.2 | 98.8 | 90.3 |
| gpt-5.2 | 98.8 | 85.9 | 100.0 | 92.4 |
| gpt-5-mini | 98.8 | 85.9 | 100.0 | 92.4 |
| Model | Acc. | Prec. | Rec. | F1 |
|---|---|---|---|---|
| gpt-5.4 | 98.0 | 96.8 | 100.0 | 98.4 |
| gpt-5.2 | 100.0 | 100.0 | 100.0 | 100.0 |
| gpt-5-mini | 100.0 | 100.0 | 100.0 | 100.0 |
| Rule | FP | FN | Prec. | Rec. |
|---|---|---|---|---|
| R06 — award within 70 days | 7 | 0 | 0.42 | 1.00 |
| R09 — end-after-publication lifecycle | 5 | 0 | 0.50 | 1.00 |
| R05 — publication/participation require award | 2 | 0 | 0.83 | 1.00 |
| Other 9 rules (each) | 0 | 0 | 1.00 | 1.00 |
With gpt-5-mini, all 14 residual errors across 1,200 decisions are false positives and zero are false negatives — no genuine violation is ever missed. Nine of twelve rules are perfect (P = R = 1.0); errors concentrate on three time-dependent rules requiring date arithmetic (the 70-day award window, the publication→award rule, and a lifecycle-timing rule) — a narrow, interpretable failure mode rather than broad unreliability.
On GDPR, where most obligations can't be verified from public ToS text, ARCCS overwhelmingly returns Insufficient Information instead of guessing — while still surfacing genuine contradictions.
Independent LLM judges agree with ARCCS verdicts in up to 96.67% of sampled cases, with substantial inter-judge agreement (Fleiss' κ = 0.69).
Results are uniformly high across three different backbones (gpt-5.4, gpt-5.2, gpt-5-mini) — performance derives from the pipeline design, not any one LLM.
ARCCS ships as a self-contained, locally-run web application. Documents are processed on the user's own machine — never uploaded to a third-party service — with the language model (hosted API or self-hosted open-weight) chosen by the user.
Prefer to explore the RPEM → CCM pipeline in code instead? Open the notebook demo in Google Colab — it installs everything for you, no local setup required.

Upload a regulatory document or pick a preloaded one (e.g. GDPR), which reuses a cached, pre-extracted requirement set.

Drag-and-drop the policy or terms-of-use document to be assessed against the extracted requirements.

Real-time logs stream pipeline progress; the final report lists every requirement with its label, confidence, and evidence.
One worked output per label, drawn from real compliance checks of the GDPR against the WhatsApp and Netflix terms-of-use documents.
"WhatsApp's Privacy Policy describes our data (including message) practices… and your rights in relation to the processing of information about you."
The document neither claims GDPR is inapplicable nor states it applies where it shouldn't — no direct conflict with the material-scope provision.
"We do not warrant that any information provided by us is accurate, complete, or useful…"
The accuracy principle is a mandatory controller obligation; the Terms explicitly disclaim it with no clause re-establishing the commitment.
Evidence from document: none located.
The Netflix excerpt contains no statements about personal-data processing or intra-EU transfer — compliance cannot be determined, so the system abstains.
"…Customer Service may best be able to assist you by using a remote access support tool through which we have full access to your computer. If you do not want us to have this access, you should not consent…"
A consent-gated remote-access clause neither clearly permits a forbidden action nor clearly satisfies the obligation — confidence falls below 0.70, so the case is deferred to a human expert.
@misc{filandrianos2026arccs,
title = {ARCCS: An Automated Regulatory Compliance Checking System},
author = {Filandrianos, Giorgos and Menezes, Jos\'e and Zerva, Chrysoula and Gianola, Alessandro},
year = {2026},
url = {https://github.com/geofila/ARCCS}
}