🕵️ Your First Anonymization¶
Detect sensitive entities and replace them with LLM-generated substitutes -- the simplest end-to-end example of Anonymizer.
📚 What you'll learn¶
- Load a CSV dataset and configure Anonymizer in a few lines
- Preview anonymized results on a small sample before committing to a full run
- Inspect entity detection and replacement with
display_record() - Process the full dataset with
run()
Tip: First time running notebooks? Start with setup instructions.
⚙️ Setup¶
- Install the notebook extra, then provide credentials for the configured external LLM providers.
create_anonymizer()starts pinned GLiNER2 locally and selects CUDA, MPS, or CPU automatically.- The default external LLM models currently use OpenRouter; its terms and privacy practices apply.
Data boundary: GLiNER2 detection runs locally in this notebook environment. LLM-assisted validation, augmentation, replacement, rewriting, repair, and evaluation use configured external hosts and may send them original or tagged input text. Do not treat this configuration as an all-local privacy boundary.
configure_logging(LoggingConfig.default())keeps logs at INFO. Switch toLoggingConfig.debug()when troubleshooting.
In [1]:
Copied!
import getpass
import os
import subprocess
import sys
package_spec = os.getenv("ANONYMIZER_NOTEBOOK_PACKAGE", "nemo-anonymizer[notebooks]")
subprocess.check_call([sys.executable, "-m", "pip", "install", "--quiet", package_spec])
import getpass
import os
import subprocess
import sys
package_spec = os.getenv("ANONYMIZER_NOTEBOOK_PACKAGE", "nemo-anonymizer[notebooks]")
subprocess.check_call([sys.executable, "-m", "pip", "install", "--quiet", package_spec])
Out[1]:
0
In [2]:
Copied!
from anonymizer.notebooks import required_api_key_environment_variables
for variable in required_api_key_environment_variables():
key = getpass.getpass(f"Enter {variable}: ").strip()
if not key:
raise RuntimeError(f"{variable} is required by the configured external model providers.")
os.environ[variable] = key
from anonymizer.notebooks import required_api_key_environment_variables
for variable in required_api_key_environment_variables():
key = getpass.getpass(f"Enter {variable}: ").strip()
if not key:
raise RuntimeError(f"{variable} is required by the configured external model providers.")
os.environ[variable] = key
In [3]:
Copied!
from anonymizer import AnonymizerConfig, AnonymizerInput, LoggingConfig, Substitute, configure_logging
from anonymizer.notebooks import create_anonymizer, stop_local_runtime
configure_logging(LoggingConfig.default())
from anonymizer import AnonymizerConfig, AnonymizerInput, LoggingConfig, Substitute, configure_logging
from anonymizer.notebooks import create_anonymizer, stop_local_runtime
configure_logging(LoggingConfig.default())
In [4]:
Copied!
anonymizer = create_anonymizer()
anonymizer = create_anonymizer()
[00:00:01] [INFO] 🔧 Anonymizer initialized with 5 model configs
[00:00:01] [INFO] |-- 🔎 detector: local-gliner2-pii
[00:00:01] [INFO] |-- ✅ validator: gpt-oss-120b
[00:00:01] [INFO] |-- 🧩 augmenter: gpt-oss-120b
GLiNER2 ready: model=fastino/gliner2-privacy-filter-PII-multi revision=59894c087cb2923b01f337d4ee72f6ff84d5bdd6 device=mps endpoint=http://127.0.0.1:52975/v1
📦 Load data and configure¶
AnonymizerInputpoints to your CSV and names the text column.data_summarygives the LLM context about the kind of text it will process.- Records up to 2,000 tokens each work with the default model configs.
AnonymizerConfigwithSubstitute()tells Anonymizer to replace detected entities with LLM-generated synthetic values for names, cities, dates, etc.
In [5]:
Copied!
input_data = AnonymizerInput(
source="https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv",
text_column="biography",
data_summary="Biographical profiles of individuals",
)
config = AnonymizerConfig(replace=Substitute())
input_data = AnonymizerInput(
source="https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv",
text_column="biography",
data_summary="Biographical profiles of individuals",
)
config = AnonymizerConfig(replace=Substitute())
👁️ Preview¶
preview()runs on a small sample so you can iterate quickly.- Always preview before processing the full dataset -- it's the fastest way to catch prompt or config issues early.
In [6]:
Copied!
preview = anonymizer.preview(config=config, data=input_data, num_records=3)
preview = anonymizer.preview(config=config, data=input_data, num_records=3)
[00:00:01] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:00:01] [INFO] 🔍 Running entity detection on 3 records
[00:00:01] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:01:41] [INFO] |-- 📋 Detection complete — 84 entities found across 3 records (0 failed) [100.1s]
[00:01:41] [INFO] |-- labels: first_name=23, occupation=6, age=5, organization_name=5, field_of_study=5, company_name=5, city=4, degree=4, university=4, last_name=3, state=3, language=3, political_view=3, place_name=2, race_ethnicity=2, street_address=2, religious_belief=2, nationality=1, date_of_birth=1, landmark=1
[00:01:41] [INFO] 🔄 Running Substitute replacement
[00:02:53] [INFO] |-- 📋 Replacement complete (0 failed) [71.7s]
[00:02:53] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
🔍 Inspect¶
display_record()shows the original text with highlighted entities, the replacement map, and the anonymized output -- all in one view.- The result dataframe has original and substituted text side-by-side.
In [7]:
Copied!
preview.display_record(0)
preview.display_record(0)
In [8]:
Copied!
preview.display_record(1)
preview.display_record(1)
In [9]:
Copied!
preview.dataframe
preview.dataframe
Out[9]:
| biography | biography_with_spans | final_entities | biography_replaced | |
|---|---|---|---|---|
| 0 | Bobby Watford, a 40‑year‑old Mexican veterinar... | <first_name>Bobby</first_name> <last_name>Watf... | {'entities': [{'id': 'first_name_0_5', 'value'... | Ethan Henderson, a 52-year-old Canadian marine... |
| 1 | Idilio Bell is a 37‑year‑old astronomer living... | <first_name>Idilio</first_name> <last_name>Bel... | {'entities': [{'id': 'first_name_0_6', 'value'... | Santiago Khan is a 42‑year‑old geophysicist li... |
| 2 | Jodi Allison, 36, lives at 204 Bluegrass in Cl... | <first_name>Jodi</first_name> <last_name>Allis... | {'entities': [{'id': 'first_name_0_4', 'value'... | Leah Keller, 42, lives at 317 Oakridge in Bent... |
🚀 Full run¶
run()processes the entire dataset with the same config you previewed.- Access the output via
result.dataframe.
In [10]:
Copied!
result = anonymizer.run(config=config, data=input_data)
print(result)
result = anonymizer.run(config=config, data=input_data)
print(result)
[00:02:54] [INFO] 📂 Loaded 25 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:02:54] [INFO] 🔍 Running entity detection on 25 records
[00:02:54] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:07:59] [INFO] |-- 📋 Detection complete — 695 entities found across 25 records (0 failed) [305.1s]
[00:07:59] [INFO] |-- labels: first_name=153, field_of_study=52, occupation=47, company_name=44, organization_name=42, city=38, university=35, last_name=26, age=26, state=25, degree=25, political_view=25, religious_belief=25, street_address=23, language=19, place_name=18, race_ethnicity=18, county=11, nationality=10, date_of_birth=9, employment_status=9, education_level=6, date=5, landmark=1, country=1, sexuality=1, postcode=1
[00:07:59] [INFO] 🔄 Running Substitute replacement
[00:11:02] [INFO] |-- 📋 Replacement complete (0 failed) [183.0s]
[00:11:02] [INFO] 🎉 Pipeline complete — 25 records processed, 0 total failures
AnonymizerResult(rows=25, columns=4, trace_columns=23, failed_records=0)
In [11]:
Copied!
result.dataframe.head()
result.dataframe.head()
Out[11]:
| biography | biography_with_spans | final_entities | biography_replaced | |
|---|---|---|---|---|
| 0 | Bobby Watford, a 40‑year‑old Mexican veterinar... | <first_name>Bobby</first_name> <last_name>Watf... | {'entities': array([{'id': 'first_name_0_5', '... | Ethan Kline, a 45‑year‑old Canadian marine bio... |
| 1 | Idilio Bell is a 37‑year‑old astronomer living... | <first_name>Idilio</first_name> <last_name>Bel... | {'entities': array([{'id': 'first_name_0_6', '... | Rafael Hawthorne is a 45‑year‑old geophysicist... |
| 2 | Jodi Allison, 36, lives at 204 Bluegrass in Cl... | <first_name>Jodi</first_name> <last_name>Allis... | {'entities': array([{'id': 'first_name_0_4', '... | Leah Harper, 42, lives at 317 Oakridge in Bent... |
| 3 | James Mills is a 69‑year‑old paramedic who liv... | <first_name>James</first_name> <last_name>Mill... | {'entities': array([{'id': 'first_name_0_5', '... | Robert Harper is a 72‑year‑old firefighter who... |
| 4 | Nancy Burton is a 21‑year‑old cashier who live... | <first_name>Nancy</first_name> <last_name>Burt... | {'entities': array([{'id': 'first_name_0_5', '... | Sophie Keller is a 27‑year‑old customer servic... |
📊 (Optional) Evaluate replacement quality¶
evaluate()is a separate, opt-in step that scores the output with LLM-as-judge metrics.- For Substitute, all four metrics run: Detection Validity, Type Fidelity, Relational Consistency, Attribute Fidelity.
- Skip it for routine runs; call it when you want LLM-side confidence on the output. Costs LLM calls per record, so try it on
previewfirst.
In [12]:
Copied!
evaluated = anonymizer.evaluate(preview)
evaluated.display_record(0)
evaluated = anonymizer.evaluate(preview)
evaluated.display_record(0)
[00:11:03] [INFO] 🧪 Running Substitute evaluation on 3 records
[00:11:03] [INFO] |-- ⚖️ Running replace judges
[00:12:31] [INFO] |-- 📋 Replace judges complete [88.3s]
[00:12:31] [INFO] 🎉 Evaluation complete — 3 records processed [88.3s]
⏭️ Next steps¶
- 🔍 Inspecting Detected Entities -- dig into what the detection pipeline found and debug quality.
- 🎯 Choosing a Replacement Strategy -- compare Redact, Annotate, Hash, and Substitute side-by-side.
- ✏️ Rewriting Biographies -- generate privacy-safe paraphrases instead of token-level replacements.
In [13]:
Copied!
stop_local_runtime()
stop_local_runtime()