🕵️ Choosing a Replacement Strategy¶
Four replace mode strategies compared side-by-side on the same data.
| Strategy | What it does |
|---|---|
| Substitute | LLM-generated contextual replacements |
| Redact | Label-based markers ([REDACTED_FIRST_NAME]) |
| Annotate | Tags entities but keeps original text |
| Hash | Deterministic hash digest |
📚 What you'll learn¶
- Compare Redact, Annotate, Hash, and Substitute on the same input
- Customize output formats with
format_template - Understand which strategy fits your use case (readability, determinism, privacy)
Tip: First time running notebooks? Start with setup instructions.
⚙️ Setup¶
- Install the notebook extra, then provide credentials for the configured external LLM providers.
create_anonymizer()starts pinned GLiNER2 locally and selects CUDA, MPS, or CPU automatically.- The default external LLM models currently use OpenRouter; its terms and privacy practices apply.
Data boundary: GLiNER2 detection runs locally in this notebook environment. LLM-assisted validation, augmentation, replacement, rewriting, repair, and evaluation use configured external hosts and may send them original or tagged input text. Do not treat this configuration as an all-local privacy boundary.
configure_logging(LoggingConfig.default())keeps logs at INFO. Switch toLoggingConfig.debug()when troubleshooting.
In [1]:
Copied!
import getpass
import os
import subprocess
import sys
package_spec = os.getenv("ANONYMIZER_NOTEBOOK_PACKAGE", "nemo-anonymizer[notebooks]")
subprocess.check_call([sys.executable, "-m", "pip", "install", "--quiet", package_spec])
import getpass
import os
import subprocess
import sys
package_spec = os.getenv("ANONYMIZER_NOTEBOOK_PACKAGE", "nemo-anonymizer[notebooks]")
subprocess.check_call([sys.executable, "-m", "pip", "install", "--quiet", package_spec])
Out[1]:
0
In [2]:
Copied!
from anonymizer.notebooks import required_api_key_environment_variables
for variable in required_api_key_environment_variables():
key = getpass.getpass(f"Enter {variable}: ").strip()
if not key:
raise RuntimeError(f"{variable} is required by the configured external model providers.")
os.environ[variable] = key
from anonymizer.notebooks import required_api_key_environment_variables
for variable in required_api_key_environment_variables():
key = getpass.getpass(f"Enter {variable}: ").strip()
if not key:
raise RuntimeError(f"{variable} is required by the configured external model providers.")
os.environ[variable] = key
In [3]:
Copied!
from anonymizer import (
Annotate,
AnonymizerConfig,
AnonymizerInput,
Hash,
LoggingConfig,
Redact,
Substitute,
configure_logging,
)
from anonymizer.notebooks import create_anonymizer, stop_local_runtime
configure_logging(LoggingConfig.default())
from anonymizer import (
Annotate,
AnonymizerConfig,
AnonymizerInput,
Hash,
LoggingConfig,
Redact,
Substitute,
configure_logging,
)
from anonymizer.notebooks import create_anonymizer, stop_local_runtime
configure_logging(LoggingConfig.default())
In [4]:
Copied!
anonymizer = create_anonymizer()
anonymizer = create_anonymizer()
[00:19:51] [INFO] 🔧 Anonymizer initialized with 5 model configs
[00:19:51] [INFO] |-- 🔎 detector: local-gliner2-pii
[00:19:51] [INFO] |-- ✅ validator: gpt-oss-120b
[00:19:51] [INFO] |-- 🧩 augmenter: gpt-oss-120b
GLiNER2 ready: model=fastino/gliner2-privacy-filter-PII-multi revision=59894c087cb2923b01f337d4ee72f6ff84d5bdd6 device=mps endpoint=http://127.0.0.1:54611/v1
📦 Input data¶
- We use the same biographies dataset throughout so each strategy is compared on identical input.
In [5]:
Copied!
input_data = AnonymizerInput(
source="https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv",
text_column="biography",
data_summary="Biographical profiles",
)
input_data = AnonymizerInput(
source="https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv",
text_column="biography",
data_summary="Biographical profiles",
)
🔄 Substitute¶
- Uses an LLM to generate contextually appropriate synthetic replacements.
- The LLM considers the full document context matching names with emails, cities to states, etc.
- Customize with
instructionsto steer the LLM's replacement choices.
In [6]:
Copied!
substitute_config = AnonymizerConfig(replace=Substitute())
substitute_preview = anonymizer.preview(
config=substitute_config,
data=input_data,
num_records=3,
)
substitute_config = AnonymizerConfig(replace=Substitute())
substitute_preview = anonymizer.preview(
config=substitute_config,
data=input_data,
num_records=3,
)
[00:19:51] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:19:51] [INFO] 🔍 Running entity detection on 3 records
[00:19:51] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:21:28] [INFO] |-- 📋 Detection complete — 84 entities found across 3 records (0 failed) [97.0s]
[00:21:28] [INFO] |-- labels: first_name=23, occupation=6, age=5, field_of_study=5, company_name=5, city=4, organization_name=4, degree=4, university=4, last_name=3, state=3, language=3, political_view=3, religious_belief=3, place_name=2, race_ethnicity=2, street_address=2, nationality=1, date_of_birth=1, landmark=1
[00:21:28] [INFO] 🔄 Running Substitute replacement
[00:22:12] [INFO] |-- 📋 Replacement complete (0 failed) [44.0s]
[00:22:12] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
In [7]:
Copied!
substitute_preview.display_record(0)
substitute_preview.display_record(0)
Custom instructions¶
- Pass
instructionsto guide the LLM -- e.g. keep replacements within a specific region, culture, or naming convention.
In [8]:
Copied!
substitute_custom_config = AnonymizerConfig(
replace=Substitute(instructions="Use only Japanese names and locations for all replacements.")
)
substitute_custom_preview = anonymizer.preview(
config=substitute_custom_config,
data=input_data,
num_records=3,
)
substitute_custom_preview.display_record(0)
substitute_custom_config = AnonymizerConfig(
replace=Substitute(instructions="Use only Japanese names and locations for all replacements.")
)
substitute_custom_preview = anonymizer.preview(
config=substitute_custom_config,
data=input_data,
num_records=3,
)
substitute_custom_preview.display_record(0)
[00:22:14] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:22:14] [INFO] 🔍 Running entity detection on 3 records
[00:22:14] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:23:33] [INFO] |-- 📋 Detection complete — 84 entities found across 3 records (0 failed) [78.6s]
[00:23:33] [INFO] |-- labels: first_name=23, occupation=6, company_name=6, age=5, field_of_study=5, city=4, organization_name=4, degree=4, university=4, last_name=3, state=3, language=3, political_view=3, place_name=2, race_ethnicity=2, street_address=2, religious_belief=2, nationality=1, date_of_birth=1, landmark=1
[00:23:33] [INFO] 🔄 Running Substitute replacement
[00:24:01] [INFO] |-- 📋 Replacement complete (0 failed) [28.5s]
[00:24:01] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
🚫 Redact¶
- Replaces each entity with a label-based marker. Default:
[REDACTED_FIRST_NAME]. - Customize with
Redact(format_template=...).
In [9]:
Copied!
redact_config = AnonymizerConfig(replace=Redact())
redact_preview = anonymizer.preview(
config=redact_config,
data=input_data,
num_records=3,
)
redact_preview.display_record(0)
redact_config = AnonymizerConfig(replace=Redact())
redact_preview = anonymizer.preview(
config=redact_config,
data=input_data,
num_records=3,
)
redact_preview.display_record(0)
[00:24:02] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:24:02] [INFO] 🔍 Running entity detection on 3 records
[00:24:02] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:25:20] [INFO] |-- 📋 Detection complete — 85 entities found across 3 records (0 failed) [77.7s]
[00:25:20] [INFO] |-- labels: first_name=23, occupation=6, age=5, university=5, field_of_study=5, organization_name=5, company_name=5, city=4, last_name=3, state=3, language=3, political_view=3, degree=3, place_name=2, race_ethnicity=2, street_address=2, religious_belief=2, nationality=1, education_level=1, date_of_birth=1, landmark=1
[00:25:20] [INFO] 🔄 Running Redact replacement
[00:25:20] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:25:20] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
Custom template¶
format_template="***"replaces every entity with the same constant.
In [10]:
Copied!
custom_config = AnonymizerConfig(replace=Redact(format_template="***"))
custom_preview = anonymizer.preview(
config=custom_config,
data=input_data,
num_records=3,
)
custom_preview.display_record(0)
custom_config = AnonymizerConfig(replace=Redact(format_template="***"))
custom_preview = anonymizer.preview(
config=custom_config,
data=input_data,
num_records=3,
)
custom_preview.display_record(0)
[00:25:21] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:25:21] [INFO] 🔍 Running entity detection on 3 records
[00:25:21] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:26:20] [INFO] |-- 📋 Detection complete — 85 entities found across 3 records (0 failed) [58.9s]
[00:26:20] [INFO] |-- labels: first_name=23, occupation=6, age=5, organization_name=5, field_of_study=5, company_name=5, city=4, university=4, last_name=3, state=3, language=3, political_view=3, degree=3, religious_belief=3, place_name=2, race_ethnicity=2, street_address=2, nationality=1, education_level=1, date_of_birth=1, landmark=1
[00:26:20] [INFO] 🔄 Running Redact replacement
[00:26:20] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:26:20] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
🏷️ Annotate¶
- Tags each entity with its label but keeps the original text visible.
Default:
<Alice, first_name>. - Customize with
format_template-- must include{text}and{label}, e.g.Annotate(format_template="<{text}-|-{label}>").
In [11]:
Copied!
annotate_config = AnonymizerConfig(replace=Annotate())
annotate_preview = anonymizer.preview(
config=annotate_config,
data=input_data,
num_records=3,
)
annotate_preview.display_record(0)
annotate_config = AnonymizerConfig(replace=Annotate())
annotate_preview = anonymizer.preview(
config=annotate_config,
data=input_data,
num_records=3,
)
annotate_preview.display_record(0)
[00:26:21] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:26:21] [INFO] 🔍 Running entity detection on 3 records
[00:26:21] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:26:58] [INFO] |-- 📋 Detection complete — 85 entities found across 3 records (0 failed) [36.7s]
[00:26:58] [INFO] |-- labels: first_name=23, occupation=6, age=5, organization_name=5, field_of_study=5, company_name=5, city=4, degree=4, university=4, last_name=3, state=3, language=3, political_view=3, religious_belief=3, place_name=2, race_ethnicity=2, street_address=2, nationality=1, date_of_birth=1, landmark=1
[00:26:58] [INFO] 🔄 Running Annotate replacement
[00:26:58] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:26:58] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
Custom template¶
- Override the default format with any string containing
{text}and{label}.
In [12]:
Copied!
annotate_custom_config = AnonymizerConfig(replace=Annotate(format_template="<{text}-|-{label}>"))
annotate_custom_preview = anonymizer.preview(
config=annotate_custom_config,
data=input_data,
num_records=3,
)
annotate_custom_preview.display_record(0)
annotate_custom_config = AnonymizerConfig(replace=Annotate(format_template="<{text}-|-{label}>"))
annotate_custom_preview = anonymizer.preview(
config=annotate_custom_config,
data=input_data,
num_records=3,
)
annotate_custom_preview.display_record(0)
[00:26:59] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:26:59] [INFO] 🔍 Running entity detection on 3 records
[00:26:59] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:27:58] [INFO] |-- 📋 Detection complete — 83 entities found across 3 records (0 failed) [59.4s]
[00:27:58] [INFO] |-- labels: first_name=23, occupation=6, age=5, field_of_study=5, company_name=5, city=4, organization_name=4, degree=4, university=4, last_name=3, state=3, language=3, political_view=3, place_name=2, race_ethnicity=2, street_address=2, religious_belief=2, nationality=1, date_of_birth=1, landmark=1
[00:27:58] [INFO] 🔄 Running Annotate replacement
[00:27:58] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:27:58] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
#️⃣ Hash¶
- Deterministic -- same input always produces the same hash.
- Customize with
format_template(must include{digest}),algorithm(sha256/sha1/md5), anddigest_length(6-64 characters).
In [13]:
Copied!
hash_config = AnonymizerConfig(replace=Hash())
hash_preview = anonymizer.preview(
config=hash_config,
data=input_data,
num_records=3,
)
hash_preview.display_record(0)
hash_config = AnonymizerConfig(replace=Hash())
hash_preview = anonymizer.preview(
config=hash_config,
data=input_data,
num_records=3,
)
hash_preview.display_record(0)
[00:27:59] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:27:59] [INFO] 🔍 Running entity detection on 3 records
[00:27:59] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:28:36] [INFO] |-- 📋 Detection complete — 83 entities found across 3 records (0 failed) [37.1s]
[00:28:36] [INFO] |-- labels: first_name=23, occupation=6, age=5, university=5, field_of_study=5, company_name=5, city=4, last_name=3, state=3, language=3, organization_name=3, political_view=3, degree=3, place_name=2, race_ethnicity=2, street_address=2, religious_belief=2, nationality=1, education_level=1, date_of_birth=1, landmark=1
[00:28:36] [INFO] 🔄 Running Hash replacement
[00:28:36] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:28:36] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
Custom template¶
- Override the algorithm, digest length, and output format.
In [14]:
Copied!
hash_custom_config = AnonymizerConfig(replace=Hash(algorithm="md5", digest_length=8, format_template="[{digest}]"))
hash_custom_preview = anonymizer.preview(
config=hash_custom_config,
data=input_data,
num_records=3,
)
hash_custom_preview.display_record(0)
hash_custom_config = AnonymizerConfig(replace=Hash(algorithm="md5", digest_length=8, format_template="[{digest}]"))
hash_custom_preview = anonymizer.preview(
config=hash_custom_config,
data=input_data,
num_records=3,
)
hash_custom_preview.display_record(0)
[00:28:37] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:28:37] [INFO] 🔍 Running entity detection on 3 records
[00:28:37] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:29:55] [INFO] |-- 📋 Detection complete — 84 entities found across 3 records (0 failed) [78.0s]
[00:29:55] [INFO] |-- labels: first_name=22, occupation=6, age=5, organization_name=5, field_of_study=5, company_name=5, city=4, degree=4, university=4, last_name=3, state=3, language=3, political_view=3, religious_belief=3, place_name=2, race_ethnicity=2, street_address=2, nationality=1, date_of_birth=1, landmark=1
[00:29:55] [INFO] 🔄 Running Hash replacement
[00:29:55] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:29:55] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
📊 (Optional) Evaluate each strategy¶
evaluate()is a separate, opt-in step that scores the output with LLM-as-judge metrics. Which metrics fire depends on the strategy:- Substitute → 4 metrics (Detection Validity + Type Fidelity + Relational Consistency + Attribute Fidelity).
- Redact / Annotate / Hash → Detection Validity only (no replacement map to score type/relational/attribute against).
- Below shows it on the Substitute preview to surface all four; the same call works on
redact_preview,annotate_preview, orhash_preview.
In [15]:
Copied!
substitute_evaluated = anonymizer.evaluate(substitute_preview)
substitute_evaluated.display_record(0)
substitute_evaluated = anonymizer.evaluate(substitute_preview)
substitute_evaluated.display_record(0)
[00:29:55] [INFO] 🧪 Running Substitute evaluation on 3 records
[00:29:55] [INFO] |-- ⚖️ Running replace judges
[00:32:47] [INFO] |-- 📋 Replace judges complete [171.3s]
[00:32:47] [INFO] 🎉 Evaluation complete — 3 records processed [171.3s]
⏭️ Next steps¶
- 🕵️ Inspecting Detected Entities -- dig into what the detection pipeline found and debug quality.
- ✏️ Rewriting Biographies -- generate privacy-safe paraphrases instead of token-level replacements.
- ⚖️ Rewriting Legal Documents -- rewrite legal text with domain-specific privacy goals.
In [16]:
Copied!
stop_local_runtime()
stop_local_runtime()