★ Fully open-source · regulation-agnostic · runs locally

ARCCS: An Automated Regulatory Compliance Checking System

An end-to-end, agentic Legal NLP system that turns raw regulatory text into atomic, traceable requirements and checks any target document against them — with retrieved evidence, calibrated confidence, and human-interpretable justification for every decision.

Giorgos Filandrianos1, José Menezes2, Chrysoula Zerva2,3, Alessandro Gianola3

1Instituto de Telecomunicações  ·  2Instituto de Telecomunicações / IST, Universidade de Lisboa  ·  3INESC-ID / IST, Universidade de Lisboa
ARCCS pipeline diagram showing RPEM and CCM

Figure 1. The ARCCS pipeline. The Regulatory Processing & Extraction Module (RPEM) turns a raw regulation into atomic, traceable requirements; the Compliance Classification Module (CCM) checks a target document against each one, returning a label, a confidence score, and cited evidence.

98.8%
rule-level violation-detection accuracy
up to 96.67%
LLM-judge evidential consistency
1,200+
individual rule checks evaluated
4
first-class labels, incl. abstention & deferral
Abstract

Decide, abstain, or defer — never guess

Regulatory compliance checking — deciding whether a target document satisfies the obligations of a regulation — requires interpreting dense legal text, identifying which provisions apply, and grounding each decision in explicit evidence. We present ARCCS, an end-to-end, automated, agentic, and regulation-agnostic Legal NLP system for compliance checking. ARCCS decomposes raw regulatory text into atomic, traceable requirements and evaluates a target document against them using retrieved evidence, confidence scores, and human-interpretable justifications.

This design decouples compliance assessment from any fixed regulatory template or predefined rule set, enabling the pipeline to operate over regulations of varying size and structure. Rather than forcing a verdict on every requirement, ARCCS assigns one of four labels — Compliant, Non-Compliant, Insufficient Information, and Human Required — making abstention and human deferral explicit outcomes.

On an EU public-procurement benchmark of more than 1,200 individual rule checks, ARCCS attains 98.8% accuracy in violation detection, and an LLM-as-a-judge evaluation finds its justifications legally and evidentially consistent in up to 96.67% of the assessed cases. ARCCS is, to our knowledge, the first fully open-source system for end-to-end regulatory compliance checking and auditable report generation.

System Architecture

Two modules, one auditable pipeline

ARCCS takes a regulatory document D and a target document T and returns a structured report stating — for every obligation in D — whether T satisfies it, with what confidence, and on the basis of which textual evidence.

Regulation (D)
→
RPEM
chunk → extract → filter
→
Requirement set ℛ
→
CCM
applicability → retrieval → label
→
Compliance report

RPEM Regulatory Processing & Extraction

Converts a raw regulation into a compact, non-redundant set of atomic requirement objects.

  • Hierarchical segmentation — splits the regulation into chunks that respect its legal structure (articles, paragraphs, sub-paragraphs), preserving local context.
  • Atomic requirement extraction — maps each chunk to a tuple (f, α, σ, M, src): legal function, regulated actor, scope, mandatory conditions, and a pointer back to the source text.
  • Filtering & consolidation — drops purely definitional text and deduplicates requirements restated across cross-references.

CCM Compliance Classification

Evaluates each applicable requirement against the target document in three steps.

  • Applicability — decides whether a requirement governs the target document at all, excluding irrelevant obligations.
  • Evidence retrieval — scores every passage and keeps only those above a relevance threshold τ, grounding the decision in retrieved text rather than parametric memory.
  • Labelling — emits a label, a confidence score κ derived from evidence coverage, and a citation-backed natural-language justification.
Decision Outcomes

Abstention and deferral are first-class labels

ARCCS never forces a verdict. With confidence thresholds θlo < θhi, every requirement resolves to one of four outcomes:

Compliant

Every mandatory condition is explicitly supported by the evidence, nothing contradicts the requirement, and confidence κ ≥ θhi.

Non‑Compliant

The evidence directly contradicts a mandatory condition, or omits one while addressing its topic, with high confidence.

Insufficient Info

The evidence references the topic but doesn't settle the mandatory conditions — principled abstention rather than a guess.

Human Required

Ambiguity, contradictory evidence, or confidence below θlo — the case is explicitly routed to a human expert.

Experiments

Evaluated under two complementary settings

A realistic setting with no deterministic gold standard (GDPR vs. real terms-of-use documents, judged by LLM-as-a-judge), and a controlled setting with deterministic ground truth (an EU public-procurement event-log benchmark).

Table 1 — LLM-as-a-judge accuracy (90 sampled checks across WhatsApp, Netflix, ChatGPT)
gpt-5.1
96.67%
gpt-5.2
90.00%
gpt-5.2-pro
90.00%
Table 2 — Inter-evaluator agreement on the binary verdict (YES/NO)
Agreementκp-value
Cohen's κ (5.1 vs 5.2)0.470.19
Cohen's κ (5.1 vs 5.2-pro)0.470.19
Cohen's κ (5.2 vs 5.2-pro)1.00< 0.001
Fleiss' κ (3 raters)0.69< 0.001

Real-world validation: a system-raised flag matches a real regulatory fine

ARCCS autonomously flagged the WhatsApp Terms of Service as conflicting with GDPR's accuracy and transparency principles. These very terms were independently found non-compliant by the Irish Data Protection Commission, which fined WhatsApp €225M under Articles 5(1)(a) and 12–14 GDPR. Similarly, in the ChatGPT Terms of Use, ARCCS flags Article 79 (effective judicial remedy) as Non-Compliant: the mandatory-arbitration / San Francisco-forum clauses conflict with EU/EEA data subjects' right to sue in their Member State of residence.

The 12 rules — Directive 2014/24/EU, Articles 4, 48, 56 & 73
#RuleDimension
1Maximum contract-amount thresholdMonetary
2Prohibition of duplicate publication of the same callDuplicate publication
3Prohibition of award before publicationTemporal / logical
4Prohibition of award before participationTemporal / logical
5Publication and participation must eventually lead to an awardTemporal / logical
6Maximum delay of 70 days between publication and awardTemporal / logical
7Prohibition of contract start without prior awardLifecycle
8Prohibition of contract end before publicationLifecycle
9Prohibition of contract end right after publication, skipping participation/awardLifecycle
10Any started contract must eventually endLifecycle
11Any ended contract must previously have startedLifecycle
12Any terminated contract must previously have been awardedLifecycle
Table 4 — Rule-level performance (100 traces × 12 rules = 1,200 decisions / model)
ModelAcc.Prec.Rec.F1
gpt-5.498.583.298.890.3
gpt-5.298.885.9100.092.4
gpt-5-mini98.885.9100.092.4
Table 5 — Case-level (per-trace) performance
ModelAcc.Prec.Rec.F1
gpt-5.498.096.8100.098.4
gpt-5.2100.0100.0100.0100.0
gpt-5-mini100.0100.0100.0100.0
Table 6 — Document-level exact match (all 12 rule decisions simultaneously correct)
gpt-5.4
83.0%
gpt-5.2
86.0%
gpt-5-mini
87.0%
Table 7 — Per-rule error breakdown (gpt-5-mini; 100 cases/rule)
RuleFPFNPrec.Rec.
R06 — award within 70 days700.421.00
R09 — end-after-publication lifecycle500.501.00
R05 — publication/participation require award200.831.00
Other 9 rules (each)001.001.00

The residual error is false-positive-driven, not miss-driven

With gpt-5-mini, all 14 residual errors across 1,200 decisions are false positives and zero are false negatives — no genuine violation is ever missed. Nine of twelve rules are perfect (P = R = 1.0); errors concentrate on three time-dependent rules requiring date arithmetic (the 70-day award window, the publication→award rule, and a lifecycle-timing rule) — a narrow, interpretable failure mode rather than broad unreliability.

01

Calibrated abstention

On GDPR, where most obligations can't be verified from public ToS text, ARCCS overwhelmingly returns Insufficient Information instead of guessing — while still surfacing genuine contradictions.

02

Evidence-grounded decisions

Independent LLM judges agree with ARCCS verdicts in up to 96.67% of sampled cases, with substantial inter-judge agreement (Fleiss' κ = 0.69).

03

Architecture, not a single model

Results are uniformly high across three different backbones (gpt-5.4, gpt-5.2, gpt-5-mini) — performance derives from the pipeline design, not any one LLM.

Demonstration

A three-step wizard, no code required

ARCCS ships as a self-contained, locally-run web application. Documents are processed on the user's own machine — never uploaded to a third-party service — with the language model (hosted API or self-hosted open-weight) chosen by the user.

Prefer to explore the RPEM → CCM pipeline in code instead? Open the notebook demo in Google Colab — it installs everything for you, no local setup required.

Step 1 — upload the regulation
1

Upload the regulation

Upload a regulatory document or pick a preloaded one (e.g. GDPR), which reuses a cached, pre-extracted requirement set.

Step 2 — upload the target document
2

Upload the target document

Drag-and-drop the policy or terms-of-use document to be assessed against the extracted requirements.

Step 3 — get the compliance report
3

Inspect the compliance report

Real-time logs stream pipeline progress; the final report lists every requirement with its label, confidence, and evidence.

Qualitative Examples

The full label space, in practice

One worked output per label, drawn from real compliance checks of the GDPR against the WhatsApp and Netflix terms-of-use documents.

Compliant  GDPR Art. 2 — Material scope  ·  conf. 0.78
"WhatsApp's Privacy Policy describes our data (including message) practices… and your rights in relation to the processing of information about you."

The document neither claims GDPR is inapplicable nor states it applies where it shouldn't — no direct conflict with the material-scope provision.

Non-Compliant  GDPR Art. 5(1)(d) — Accuracy principle  ·  conf. 0.62
"We do not warrant that any information provided by us is accurate, complete, or useful…"

The accuracy principle is a mandatory controller obligation; the Terms explicitly disclaim it with no clause re-establishing the commitment.

Insufficient Info  GDPR Art. 1 — Subject matter & objectives
Evidence from document: none located.

The Netflix excerpt contains no statements about personal-data processing or intra-EU transfer — compliance cannot be determined, so the system abstains.

Human Required  Regulation (EU) 2016/679 — General obligations
"…Customer Service may best be able to assist you by using a remote access support tool through which we have full access to your computer. If you do not want us to have this access, you should not consent…"

A consent-gated remote-access clause neither clearly permits a forbidden action nor clearly satisfies the obligation — confidence falls below 0.70, so the case is deferred to a human expert.

Citation

If you use ARCCS, please cite

@misc{filandrianos2026arccs,
  title  = {ARCCS: An Automated Regulatory Compliance Checking System},
  author = {Filandrianos, Giorgos and Menezes, Jos\'e and Zerva, Chrysoula and Gianola, Alessandro},
  year   = {2026},
  url    = {https://github.com/geofila/ARCCS}
}