🕵️ Inspecting Detected Entities¶
Dig into the entity detection pipeline output -- what was detected, what the LLM validator kept or dropped, and where entities appear in the text.
This notebook is for users who need to debug detection quality, tune labels and/or thresholds, or investigate downstream replacement or rewriting results.
We use Annotate mode because it preserves the original text while tagging each entity with its label, making it ideal for reviewing detection quality.
Privacy warning:
Annotatedoes not anonymize the text. Sensitive values remain in the output, so use it only for inspection -- not as a privacy-safe production strategy.
📚 What you'll learn¶
- Run the detection pipeline and inspect its output using Annotate mode
- View tagged text with entities marked inline
- Break down detected entities by label, source, and unique value
- Identify and triage failed records
Tip: First time running notebooks? Start with setup instructions.
⚙️ Setup¶
- Install the notebook extra, then provide credentials for the configured external LLM providers.
create_anonymizer()starts pinned GLiNER2 locally and selects CUDA, MPS, or CPU automatically.- The default external LLM models currently use OpenRouter; its terms and privacy practices apply.
Data boundary: GLiNER2 detection runs locally in this notebook environment. LLM-assisted validation, augmentation, replacement, rewriting, repair, and evaluation use configured external hosts and may send them original or tagged input text. Do not treat this configuration as an all-local privacy boundary.
configure_logging(LoggingConfig.default())keeps logs at INFO. Switch toLoggingConfig.debug()when troubleshooting.
import getpass
import os
import subprocess
import sys
from collections import Counter
package_spec = os.getenv("ANONYMIZER_NOTEBOOK_PACKAGE", "nemo-anonymizer[notebooks]")
subprocess.check_call([sys.executable, "-m", "pip", "install", "--quiet", package_spec])
0
import pandas as pd
from anonymizer.notebooks import required_api_key_environment_variables
for variable in required_api_key_environment_variables():
key = getpass.getpass(f"Enter {variable}: ").strip()
if not key:
raise RuntimeError(f"{variable} is required by the configured external model providers.")
os.environ[variable] = key
from anonymizer import Annotate, AnonymizerConfig, AnonymizerInput, LoggingConfig, configure_logging
from anonymizer.notebooks import create_anonymizer, stop_local_runtime
configure_logging(LoggingConfig.default())
anonymizer = create_anonymizer()
[00:13:32] [INFO] 🔧 Anonymizer initialized with 5 model configs
[00:13:32] [INFO] |-- 🔎 detector: local-gliner2-pii
[00:13:32] [INFO] |-- ✅ validator: gpt-oss-120b
[00:13:32] [INFO] |-- 🧩 augmenter: gpt-oss-120b
GLiNER2 ready: model=fastino/gliner2-privacy-filter-PII-multi revision=59894c087cb2923b01f337d4ee72f6ff84d5bdd6 device=mps endpoint=http://127.0.0.1:54053/v1
👁️ Preview¶
- Detection runs as part of any strategy.
Annotatekeeps original text visible alongside entity labels -- ideal for debugging. trace_dataframeexposes every internal pipeline column; that's what we explore below.
config = AnonymizerConfig(replace=Annotate())
input_data = AnonymizerInput(
source="https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv",
text_column="biography",
data_summary="Biographical profiles",
)
result = anonymizer.preview(
config=config,
data=input_data,
num_records=3,
)
[00:13:33] [INFO] 👀 Preview mode: 📂 Loaded 3 records from https://raw.githubusercontent.com/NVIDIA-NeMo/Anonymizer/refs/heads/main/docs/data/NVIDIA_synthetic_biographies.csv (column: 'biography')
[00:13:33] [INFO] 🔍 Running entity detection on 3 records
[00:13:33] [INFO] detection labels in scope: (default: 65 labels; see anonymizer.DEFAULT_ENTITY_LABELS for list)
[00:16:46] [INFO] |-- 📋 Detection complete — 84 entities found across 3 records (0 failed) [192.8s]
[00:16:46] [INFO] |-- labels: first_name=23, occupation=6, age=5, field_of_study=5, company_name=5, city=4, organization_name=4, university=4, last_name=3, state=3, language=3, political_view=3, degree=3, religious_belief=3, place_name=2, race_ethnicity=2, street_address=2, nationality=1, education_level=1, date_of_birth=1, landmark=1
[00:16:46] [INFO] 🔄 Running Annotate replacement
[00:16:46] [INFO] |-- 📋 Replacement complete (0 failed) [0.0s]
[00:16:46] [INFO] 🎉 Pipeline complete — 3 records processed, 0 total failures
🔍 Inspect¶
display_record()renders an interactive view with entity highlights.
result.display_record(0)
📋 Columns¶
result.dataframe["final_entities"]is the stable, public entity output.trace_dataframecontains internal pipeline columns for deeper debugging; those underscore-prefixed columns may change between releases.
trace_df = result.trace_dataframe
final_entities = result.dataframe["final_entities"]
print(f"Records: {len(trace_df)}")
print(f"Columns: {list(trace_df.columns)}")
Records: 3 Columns: ['biography', '_anonymizer_record_id', '_raw_detected_entities', '_seed_entities', '_tag_notation', '_seed_validation_candidates', '_seed_tagged_text', '_validated_entities', '_seed_entities_json', '_initial_tagged_text', '_validated_seed_entities', '_augmented_entities', '_merged_entities', '_merged_tagged_text', '_validation_candidates', '_detected_entities', 'biography_with_spans', 'final_entities', '_entities_by_value', '_replacement_map', 'biography_replaced', '_replacement_application']
🎯 Detected entities¶
- Final entity list after validation. Each entity has
value,label, positions,score, andsource(detector / augmenter / name_split / propagation).
row_idx = 0
raw = final_entities.iloc[row_idx]
entities = raw["entities"] if isinstance(raw, dict) else raw
print(f"Record {row_idx}: {len(entities)} entities detected\n")
entity_df = pd.DataFrame(entities)
if not entity_df.empty:
cols = [c for c in ["value", "label", "start_position", "end_position", "source"] if c in entity_df.columns]
print(entity_df[cols].to_string())
Record 0: 23 entities detected
value label start_position end_position source
0 Bobby first_name 0 5 detector
1 Watford last_name 6 13 detector
2 40‑year‑old age 17 28 detector
3 Mexican nationality 29 36 detector
4 veterinarian occupation 37 49 detector
5 Denver city 60 66 detector
6 Colorado state 68 76 detector
7 Jefferson High organization_name 180 194 augmenter
8 DVM education_level 210 213 detector
9 University of Colorado Boulder university 221 251 detector
10 wildlife health field_of_study 297 312 detector
11 English language 324 331 detector
12 Bobby first_name 333 338 detector
13 animal care field_of_study 512 523 detector
14 Bobby first_name 556 561 detector
15 VCA Animal Hospital organization_name 576 595 detector
16 Colorado Veterinary Clinic organization_name 613 639 detector
17 Christian Democrat political_view 707 725 detector
18 Maya first_name 798 802 detector
19 Aria first_name 836 840 detector
20 Leo first_name 845 848 detector
21 Bobby first_name 870 875 detector
22 Rockies place_name 894 901 detector
🏷️ Labels¶
- Entity label distribution across all records -- which types are most common.
label_counts = Counter()
for raw in final_entities:
entity_list = raw["entities"] if isinstance(raw, dict) else raw
for entity in entity_list:
label_counts[entity["label"]] += 1
for label, count in label_counts.most_common():
print(f" {label}: {count}")
first_name: 23 occupation: 6 age: 5 field_of_study: 5 company_name: 5 city: 4 organization_name: 4 university: 4 last_name: 3 state: 3 language: 3 political_view: 3 degree: 3 religious_belief: 3 place_name: 2 race_ethnicity: 2 street_address: 2 nationality: 1 education_level: 1 date_of_birth: 1 landmark: 1
📡 Sources¶
- Where each entity came from in the pipeline:
detector-- GLiNER NERaugmenter-- LLM-added (missed by GLiNER)validator-- LLM decision step over detector-seed entities (keep/reclass/drop); does not emit a separate source valuename_split-- derived from splitting full namespropagation-- expanded from validated entities to all text occurrences
source_counts = Counter()
for raw in final_entities:
entity_list = raw["entities"] if isinstance(raw, dict) else raw
for entity in entity_list:
source_counts[entity.get("source", "unknown")] += 1
for source, count in source_counts.most_common():
print(f" {source}: {count}")
detector: 81 augmenter: 3
📊 By value¶
- Entities grouped by unique value -- this is what drives consistent replacement downstream (same name always maps to the same substitute).
row_idx = 0
raw_bv = trace_df.loc[row_idx, "_entities_by_value"]
by_value = raw_bv["entities_by_value"] if isinstance(raw_bv, dict) else raw_bv
print(f"Record {row_idx}: {len(by_value)} unique entity values\n")
for entry in by_value:
print(f" {entry['value']!r} -> labels: {entry['labels']}")
Record 0: 20 unique entity values '40‑year‑old' -> labels: ['age'] 'Aria' -> labels: ['first_name'] 'Bobby' -> labels: ['first_name'] 'Christian Democrat' -> labels: ['political_view'] 'Colorado' -> labels: ['state'] 'Colorado Veterinary Clinic' -> labels: ['organization_name'] 'DVM' -> labels: ['education_level'] 'Denver' -> labels: ['city'] 'English' -> labels: ['language'] 'Jefferson High' -> labels: ['organization_name'] 'Leo' -> labels: ['first_name'] 'Maya' -> labels: ['first_name'] 'Mexican' -> labels: ['nationality'] 'Rockies' -> labels: ['place_name'] 'University of Colorado Boulder' -> labels: ['university'] 'VCA Animal Hospital' -> labels: ['organization_name'] 'Watford' -> labels: ['last_name'] 'animal care' -> labels: ['field_of_study'] 'veterinarian' -> labels: ['occupation'] 'wildlife health' -> labels: ['field_of_study']
❌ Failures¶
- Records dropped during detection (LLM timeout, parse error, etc.).
- Check this to understand data loss in your pipeline.
if result.failed_records:
for fr in result.failed_records:
print(f" record_id={fr.record_id}, step={fr.step}, reason={fr.reason}")
else:
print("No failed records.")
No failed records.
📊 (Optional) Score the detections with an LLM judge¶
evaluate()is a separate, opt-in step that runs LLM-as-judge metrics on the output.- This notebook uses Annotate, so only Detection Validity runs — it flags entities the detector got wrong (false positives, mislabels, boundary errors). Substitute would also enable Type Fidelity, Relational Consistency, and Attribute Fidelity.
evaluated = anonymizer.evaluate(result)
evaluated.display_record(0)
[00:16:46] [INFO] 🧪 Running Annotate evaluation on 3 records
[00:16:46] [INFO] |-- ⚖️ Running replace judges
[00:18:50] [INFO] |-- 📋 Replace judges complete [123.3s]
[00:18:50] [INFO] 🎉 Evaluation complete — 3 records processed [123.3s]
⏭️ Next steps¶
- 🕵️ Your First Anonymization -- the simplest end-to-end replace workflow if you haven't run it yet.
- 🎯 Choosing a Replacement Strategy -- compare Redact, Annotate, Hash, and Substitute side-by-side.
- ✏️ Rewriting Biographies -- generate privacy-safe paraphrases instead of token-level replacements.
stop_local_runtime()