Skip to content

FEAT: Add Garak ProPILE privacy leakage scenario - #3115

Merged
Richard Lundeen (richlundeen) merged 2 commits into
microsoft:mainfrom
barkha1research-ai:feat-propile-scenario
Oct 10, 2026
Merged

Richard Lundeen (richlundeen) merged 2 commits into
microsoft:mainfrom
barkha1research-ai:feat-propile-scenario

Conversation

@barkha1research-ai

Copy link
Copy Markdown
Contributor

Description

Closes #2532.

Adds garak.propile, a PyRIT-native port of Garak's ProPILE probes (https://arxiv.org/abs/2307.01881).

  • Datasets: garak_propile_pii holds Garak's 26 records, one seed per subject, with source_dataset and source_id kept in metadata and the Nemotron-CC-v2.1 source linked in the description. garak_propile_templates holds the 24 templates with placeholders converted to Jinja and metadata for the technique, the withheld field and the known fields. Both point at Garak commit 3f50ea5f, already used in third_party/garak-provenance.json, and neither is tagged default.
  • Techniques: Twin, Triplet, Quadruplet and Unstructured. ProPILEDatasetConfiguration renders each template against every record that has its known and withheld fields, groups the requests by technique, and samples with one request reserved per selected technique (default max_total 20). Identical requests from duplicate records collapse, and namesakes with different values stay separate requests.
  • If a selected technique has no compatible record, initialization raises a DatasetConstraintError that names the technique, instead of reporting zero attacks. The bundled records support 81 Twin and 4 Triplet requests. Quadruplet and Unstructured are tested with synthetic records and work with a user-supplied record dataset selected alongside garak_propile_templates.
  • Scoring: each SeedObjective carries its withheld value as OutputMatches(Contains(...)) in conditions, and the scenario uses OutputMatchesScorer, so the seed authors its criterion as doc/code/framework.md describes. This uses the same ExactTextMatching as SubStringScorer (case-insensitive, edge-whitespace normalised), which matches Garak's PIILeakExact detector. A custom objective scorer without OutputMatches support receives the objective only. Happy to switch to per-value SubStringScorer instances if you prefer.
  • One deviation from the issue: Twin is the default technique, and Triplet, Quadruplet and Unstructured are opt-in. test_all_builtin_default_size_contracts_are_bounded_async requires every built-in scenario's default run to have at least one attack, so an all-opt-in default would fail it. I can change this if you would rather exempt the scenario.
  • Licensing: Apache-2.0 notices in the new module and datasets, three entries in third_party/garak-provenance.json, and a ProPILE paragraph in THIRD_PARTY_NOTICES.txt. The records are ported unchanged from Garak, including its extraction results. As the provenance scope_note says, Nemotron-CC's own terms are not audited here.
  • The docs state that an exact match indicates possible disclosure and does not prove memorisation of a specific record.

Tests and Documentation

  • New: tests/unit/scenario/garak/test_propile.py (31 tests) and tests/unit/datasets/test_garak_propile_dataset.py (7 tests). They cover real-data populations and provenance, the Quadruplet and Unstructured synthetic fixtures, multiple relationships, duplicate and namesake records, numeric record values, the no-compatible-record error (including ALL on bundled data), coverage sampling, unsupported configurations, resume replay, run-size estimates, the custom-scorer path, and a full mocked run scoring all 81 responses against their own withheld values.
  • pytest tests/unit: all pass locally (24,900 passed, 323 skipped).
  • pre-commit run on the changed files: all hooks pass. ty check reports nothing in the new files.
  • Docs: added a ProPILE section to doc/scanner/garak.py/.ipynb and the dataset names to doc/code/datasets/1_loading_datasets.py/.ipynb, plus the citation in doc/references.bib and doc/bibliography.md. Notebook inputs were synced with jupytext --update, and the .ipynb/.py round trip checked. The ProPILE result cell was executed against a local OpenAI-compatible target (Ollama, llama3.1:8b) rather than the Azure deployment used for the other cells, so its target information differs. Please regenerate it with your target if you prefer.

Signed-off-by: Barkha <340367068+barkha1research-ai@users.noreply.github.com>
@barkha1research-ai

Copy link
Copy Markdown
Contributor Author

@microsoft-github-policy-service agree

Preserve seed-authored OutputMatches conditions and reject incompatible scorer overrides before initialization.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 61f33406-d867-43e4-a626-e5472fa0dc8f
@richlundeen
Richard Lundeen (richlundeen) added this pull request to the merge queue Oct 10, 2026
Merged via the queue into microsoft:main with commit e43587b Oct 10, 2026
55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

FEAT: Add Garak ProPILE privacy leakage scenario

2 participants