🕵️ Rewriting Legal Documents¶
Rewriting legal text (TAB dataset) with a domain-specific privacy goal and custom entity labels tailored for legal proceedings.
📚 What you'll learn¶
- Define domain-specific entity labels for legal text (case numbers, court names, etc.)
- Configure rewrite mode with legal-specific privacy goals
- Preview and run on court decision documents
- Triage flagged records with
needs_human_review
Tip: First time running notebooks? Start with setup instructions.
⚙️ Setup¶
- Install the notebook extra, then provide credentials for the configured external LLM providers.
create_anonymizer()starts pinned GLiNER2 locally and selects CUDA, MPS, or CPU automatically.- The default external LLM models currently use OpenRouter; its terms and privacy practices apply.
Data boundary: GLiNER2 detection runs locally in this notebook environment. LLM-assisted validation, augmentation, replacement, rewriting, repair, and evaluation use configured external hosts and may send them original or tagged input text. Do not treat this configuration as an all-local privacy boundary.
configure_logging(LoggingConfig.default())keeps logs at INFO. Switch toLoggingConfig.debug()when troubleshooting.
import getpass
import os
import subprocess
import sys
package_spec = os.getenv("ANONYMIZER_NOTEBOOK_PACKAGE", "nemo-anonymizer[notebooks]")
subprocess.check_call([sys.executable, "-m", "pip", "install", "--quiet", package_spec])
0
from anonymizer.notebooks import required_api_key_environment_variables
for variable in required_api_key_environment_variables():
key = getpass.getpass(f"Enter {variable}: ").strip()
if not key:
raise RuntimeError(f"{variable} is required by the configured external model providers.")
os.environ[variable] = key
from anonymizer import (
AnonymizerConfig,
AnonymizerInput,
Detect,
LoggingConfig,
PrivacyGoal,
Rewrite,
configure_logging,
)
from anonymizer.notebooks import create_anonymizer, stop_local_runtime
configure_logging(LoggingConfig.default())
anonymizer = create_anonymizer()
[01:04:34] [INFO] 🔧 Anonymizer initialized with 5 model configs
[01:04:34] [INFO] |-- 🔎 detector: local-gliner2-pii
[01:04:34] [INFO] |-- ✅ validator: gpt-oss-120b
[01:04:34] [INFO] |-- 🧩 augmenter: gpt-oss-120b
GLiNER2 ready: model=fastino/gliner2-privacy-filter-PII-multi revision=59894c087cb2923b01f337d4ee72f6ff84d5bdd6 device=mps endpoint=http://127.0.0.1:57432/v1
📦 Input data¶
- TAB (Text Anonymization Benchmark) legal documents -- court decisions containing names, dates, case numbers, and other legal identifiers.
LEGAL_ENTITY_LABELSdefines the domain-specific entity types to detect. This replaces the default label set with one tailored to legal text.
LEGAL_ENTITY_LABELS = [
"first_name",
"last_name",
"court_name",
"organization_name",
"company_name",
"prison_detention_facility",
"street_address",
"city",
"state",
"country",
"date",
"date_time",
"time",
"date_of_birth",
"age",
"email",
"phone_number",
"ssn",
"unique_id",
"legal_role",
"case_number",
"application_number",
"monetary_amount",
"sentence_duration",
"nationality",
]
input_data = AnonymizerInput(
source="https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/TAB_legal_sample25.csv",
text_column="text",
data_summary="Legal court decisions containing personal identifiers, case numbers, and institutional references",
)
🎛️ Configure¶
Detect(entity_labels=...)overrides the default entity set with legal-specific labels. The explicit list is a strict allowlist for both detection and LLM augmentation: labels not included here are filtered out, so include every entity type you need.PrivacyGoaltells the rewriter what to protect (identifiers, case numbers, institutional references) and what to preserve (legal reasoning, statutory references, ruling structure).
config = AnonymizerConfig(
detect=Detect(
entity_labels=LEGAL_ENTITY_LABELS,
),
rewrite=Rewrite(
privacy_goal=PrivacyGoal(
protect="All personal identifiers, case numbers, court names, and institutional references that could identify parties",
preserve="Legal reasoning, procedural facts, statutory references, and the structure of the ruling",
),
risk_tolerance="minimal",
max_repair_iterations=3,
),
)
👁️ Preview¶
- Preview on a few records to check that legal entities are detected and the rewrite preserves the ruling's structure.
preview = anonymizer.preview(
config=config,
data=input_data,
num_records=3,
)
preview.display_record(0)
[01:04:34] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/TAB_legal_sample25.csv (column: 'text')
[01:04:34] [INFO] 🔍 Running entity detection on 3 records
[01:04:34] [INFO] detection labels in scope: ['age', 'application_number', 'case_number', 'city', 'company_name', 'country', 'court_name', 'date', 'date_of_birth', 'date_time', 'email', 'first_name', 'last_name', 'legal_role', 'monetary_amount', 'nationality', 'organization_name', 'phone_number', 'prison_detention_facility', 'sentence_duration', 'ssn', 'state', 'street_address', 'time', 'unique_id']
[01:07:37] [INFO] |-- 📋 Detection complete — 202 entities found across 3 records (0 failed) [182.6s]
[01:07:37] [INFO] |-- labels: legal_role=58, date=49, court_name=31, organization_name=24, last_name=8, application_number=5, country=5, first_name=5, city=5, date_of_birth=4, nationality=3, sentence_duration=2, monetary_amount=2, case_number=1
[01:07:37] [INFO] ✏️ Running rewrite pipeline
[01:21:56] [INFO] Evaluate-repair loop iteration 0: 2/3 rows need repair
[01:22:52] [INFO] Evaluate-repair loop: all rows pass at iteration 1
[01:22:52] [INFO] |-- 📋 Rewrite complete (0 failed) [915.0s]
[01:22:52] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
preview.display_record(1)
How to interpret leakage: Leakage is measured against the sensitivity disposition. Details marked
leave_as_ismay remain without increasingleakage_mass. If an output retains something you expected the privacy goal to protect, inspect the Entity Disposition table.
🚀 Full run¶
result.dataframehas user-facing columns: rewritten text, scores, and the review flag.- This notebook uses
risk_tolerance="minimal", which applies stricter repair and review thresholds than notebook 04.
result = anonymizer.run(config=config, data=input_data)
result.dataframe.head()
[01:22:53] [INFO] 📂 Loaded 25 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/TAB_legal_sample25.csv (column: 'text')
[01:22:53] [INFO] 🔍 Running entity detection on 25 records
[01:22:53] [INFO] detection labels in scope: ['age', 'application_number', 'case_number', 'city', 'company_name', 'country', 'court_name', 'date', 'date_of_birth', 'date_time', 'email', 'first_name', 'last_name', 'legal_role', 'monetary_amount', 'nationality', 'organization_name', 'phone_number', 'prison_detention_facility', 'sentence_duration', 'ssn', 'state', 'street_address', 'time', 'unique_id']
[01:31:46] [INFO] |-- 📋 Detection complete — 1843 entities found across 25 records (0 failed) [532.2s]
[01:31:46] [INFO] |-- labels: legal_role=550, date=399, organization_name=281, court_name=239, last_name=70, first_name=69, city=43, country=42, nationality=33, application_number=29, date_of_birth=27, sentence_duration=23, prison_detention_facility=16, monetary_amount=12, case_number=4, company_name=2, unique_id=1, age=1, time=1, state=1
[01:31:46] [INFO] ✏️ Running rewrite pipeline
[01:33:42] [WARNING] Replacement map repaired synthetic-original collision entries for record 46e1a5989f27571893c1167cf94e9bb1; repaired=1 (repaired_by_label={'legal_role': 1})
[01:36:17] [WARNING] Required rewrite replacement was unavailable; marking rewritten text unavailable.
[01:44:44] [WARNING] Evaluator returned malformed privacy answer set; missing=[36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46] duplicate=[] extra=[]. Applying conservative normalization.
[01:44:45] [INFO] Evaluate-repair loop iteration 0: 20/25 rows need repair
[01:51:42] [INFO] Evaluate-repair loop iteration 1: 11/25 rows need repair
[01:54:36] [INFO] Evaluate-repair loop iteration 2: 9/25 rows need repair
[01:58:23] [INFO] |-- 📋 Rewrite complete (0 failed) [1597.2s]
[01:58:23] [INFO] 🎉 Pipeline complete — 25 records processed, 0 total failures
| text | text_rewritten | utility_score | leakage_mass | weighted_leakage_rate | any_high_leaked | needs_human_review | |
|---|---|---|---|---|---|---|---|
| 0 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 0.988235 | 0.0 | 0.0 | False | False |
| 1 | PROCEDURE The case originated in an applicati... | The case originated in an application (no. 415... | 0.975 | 0.0 | 0.0 | False | False |
| 2 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 0.95 | 4.63 | 0.2315 | True | True |
| 3 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 0.935714 | 1.9 | 0.082609 | True | True |
| 4 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 1.0 | 2.0 | 0.166667 | True | True |
result.dataframe[["text_rewritten", "utility_score", "leakage_mass", "needs_human_review"]].head()
| text_rewritten | utility_score | leakage_mass | needs_human_review | |
|---|---|---|---|---|
| 0 | PROCEDURE The case originated in an applicati... | 0.988235 | 0.0 | False |
| 1 | The case originated in an application (no. 415... | 0.975 | 0.0 | False |
| 2 | PROCEDURE The case originated in an applicati... | 0.95 | 4.63 | True |
| 3 | PROCEDURE The case originated in an applicati... | 0.935714 | 1.9 | True |
| 4 | PROCEDURE The case originated in an applicati... | 1.0 | 2.0 | True |
🚩 Filter by review flag¶
- Records where automated metrics exceed thresholds are flagged for manual review.
- The repair loop stops after
max_repair_iterations; records that still need repair remain flagged for human review but are not pipeline failures. - Use this to prioritize human attention on the records that need it most.
- See Working with flagged records for guidance on diagnosing and resolving flagged records.
df = result.dataframe
flagged = df[df["needs_human_review"] == True] # noqa: E712
print(f"{len(flagged)} of {len(df)} records flagged for human review")
flagged.head()
10 of 25 records flagged for human review
| text | text_rewritten | utility_score | leakage_mass | weighted_leakage_rate | any_high_leaked | needs_human_review | |
|---|---|---|---|---|---|---|---|
| 2 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 0.95 | 4.63 | 0.2315 | True | True |
| 3 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 0.935714 | 1.9 | 0.082609 | True | True |
| 4 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 1.0 | 2.0 | 0.166667 | True | True |
| 6 | PROCEDURE The case originated in an applicati... | <NA> | 0.0 | 0.0 | 0.0 | False | True |
| 11 | PROCEDURE The case originated in an applicati... | PROCEDURE The case originated in an applicati... | 0.84 | 9.42 | 0.359542 | True | True |
🔬 Evaluate (optional)¶
Call evaluate() to run LLM-as-judge scoring on the rewrite result — detection validity and three quality rubrics (privacy, quality, style).
Evaluation makes additional LLM calls per record. For larger datasets, evaluate
a preview first; this tutorial evaluates all 25 rows to demonstrate the complete workflow.
This holistic judge is independent of pipeline leakage scoring, so their assessments may differ.
See Evaluation for details.
evaluated = anonymizer.evaluate(result)
[01:58:23] [INFO] 🧪 Running rewrite evaluation on 25 records
[01:58:23] [INFO] |-- ⚖️ Running rewrite judges
[02:00:25] [INFO] |-- 📋 Rewrite judges complete [121.5s]
[02:00:25] [INFO] |-- 🔎 Running entity coverage
[02:04:34] [INFO] |-- 📋 Entity coverage complete [249.5s]
[02:04:34] [INFO] 🎉 Evaluation complete — 25 records processed [371.0s]
evaluated.display_record(0)
⏭️ Next steps¶
- 📊 Evaluation -- learn about the detection validity and rewrite quality judges in detail.
- 🔍 Inspecting Detected Entities -- debug what the detection pipeline found before rewriting.
- Try it on your own data! Swap in your CSV, define entity labels for your
domain, and set a
PrivacyGoalthat fits -- you've got all the building blocks.
stop_local_runtime()