Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
LLM-Based Adversarial Text Anonymization: Evaluating Attacker Architectures via Explicit and Implicit Signal Reasoning
Stockholm University, Faculty of Social Sciences, Department of Computer and Systems Sciences.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Introduction: Large language models can infer personal attributes such as age, occupation, income, and location from everyday online text, even when explicit identifiers are removed. These inferences rely not only on directly stated information but also on implicit signals embedded in writing style, vocabulary, and broader language-use patterns. This creates privacy risks for individuals who share narrative text online and highlights the limitations of existing anonymization methods.

Research Question: This thesis investigates to what extent separating an adversarial attacker into signal-specialized components affects post-anonymization privacy protection compared with a unified attacker.

Method: This thesis adopts an exploratory experimental methodology based on a controlled comparative evaluation of adversarial anonymization architectures. Four adversarial anonymization pipeline architectures were implemented and evaluated: an explicit-only baseline, a combined single-prompt attacker, a parallel dual-attacker architecture, and a sequential coordinated dual-attacker architecture. All configurations used GPT-4o as attacker and anonymizer and were evaluated over two anonymization rounds on 100 synthetic profiles from the SynthPAI benchmark, using five metrics: Top-1 adversarial accuracy, Top-3 adversarial accuracy, evidence rate, average attacker certainty, and combined text utility.

Results: All four configurations converge to a similar post-anonymization Top-3 adversarial accuracy range of approximately 0.24–0.27. Increasing attacker specialization and coordination does not substantially improve anonymization performance relative to the explicit-only baseline. However, the dual-attacker architectures reveal a consistent asymmetry: explicit textual evidence decreases after anonymization, whereas implicit writing-style signals persist more strongly across configurations. McNemar’s test confirmed that no pairwise configuration comparison reached statistical significance.

Discussion: The findings suggest that the primary limitation of the evaluated adversarial anonymization framework lies less in attacker awareness than in the anonymizer’s limited ability to suppress distributed stylistic signals through localized rewriting. Future improvements in LLM-based anonymization may therefore depend more on redesigning anonymizers to address broader language patterns, through style-transfer methods or controllable generation, than on increasing the complexity of attackers.

Place, publisher, year, edition, pages
2026.
Keywords [en]
Large Language Models, Adversarial Text Anonymization, Privacy-Preserving NLP, Implicit Signals, Prompt-Based Reasoning, GPT-4o, Attribute Inference
National Category
Artificial Intelligence
Identifiers
URN: urn:nbn:se:su:diva-257319OAI: oai:DiVA.org:su-257319DiVA, id: diva2:2079926
Available from: 2026-06-25 Created: 2026-06-25

Open Access in DiVA

fulltext(4313 kB)74 downloads
File information
File name FULLTEXT01.pdfFile size 4313 kBChecksum SHA-512
adbc329e747aaa5c54eac5be2af0218f34a99b919851587f3daf4b27bd47672046ebbc89a9f124de3dd6dcbf93b768f17d668fa4907ff26400226401b6976e53
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Ebrahimitofighi, Negin
By organisation
Department of Computer and Systems Sciences
Artificial Intelligence

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 196 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf