Lucca Pfründer

Methodology and Statistics
Tilburg School of Social and Behavioral Sciences (TSB)
Tilburg University

Email
Website

Project
Adversarial machine learning on text data for psychological inference

Text data offer exciting potential for psychological research as they can provide rich accounts of people’s experiences or internal states. Recent advances in computational methods originating from natural language processing (NLP) and machine learning (ML) promise to investigate such data systematically (e.g., in text classification, which could inform measurement in psychology) (Mikolov et al., 2013; Vaswani et al., 2017). For example, automated deception classifiers that are trained on verbal deception data often outperform humans in distinguishing truthful and deceptive text, showing classification accuracies of around 70–80% (Bond & DePaulo, 2006; Constâncio et al., 2023). While these are exciting developments, the models underlying automated text classification bring about their own caveats and, as we argue, opportunities: adversarial attacks.

Adversarial attacks are minute, yet targeted perturbations of input data (e.g., paraphrases) that can trick a model into misclassification (Goodfellow et al., 2014; Szegedy et al., 2013). They are a test of robustness: if a model generalizes well, small changes that leave the meaning of the text intact should not affect the prediction. If it does, however, this suggests that the model might rely on patterns that are predictive in-sample but fail to generalize robustly (i.e., overfitting). Generally, these attacks can occur in white-box settings, where the attacker has access to the model’s parameters, or in black-box settings, with access only to model outputs (Ebrahimi et al., 2018). In text classification, they usually target characters, words, or phrases, are often (not always) iterative, and can involve live model feedback (Alzantot et al., 2018; Iyyer et al., 2018).

While adversarial attacks pose a threat to the validity of AI classifiers, they also offer a unique way to measure human or automated behaviour, revealing, i.e., how attackers reason about a classifier’s workings (Kleinberg et al., 2025; Mozes et al., 2021, 2022). For example, when a human attacker iteratively perturbs a text through paraphrasing, they receive live feedback on how each modification affects the model’s classification, making it possible to trace which changes they expect to matter and, in turn, which linguistic features they infer the classifier relies on. While this logic can be applied more broadly and to other constructs, in deception research, this may reveal cues that people implicitly believe to predict deception, which explicit measures often fail to capture accurately (DePaulo et al., 2003). This project fits IOPS well because of its psychometric nature. It develops and explores a novel paradigm to measure human and automated behavior. The approach further bridges disciplines and integrates new computational measurements as instruments for psychological and social-scientific research.

Supervisors
Prof. Dr. Bruno Verschuere
Dr. Bennett Kleinberg

Financed by
Department of Methodology and Statistics, Tilburg University

Period
15 january 2026 – 15 January 2032