Improving Human Judgment Estimation using Reinforcement Learning with Verifiable Rewards
cental | Louvain-la-Neuve
PhD student : Sebastian Loftus (website)
Supervisor : Marie-Catherine de Marneffe (website)
Funding : FSR Project 2024 « Towards robust Natural Language Understanding: Taking into account cognitive bias that leads to multiple interpretations »
Start : September 2025
Summary
Large Language Models (LLMs) currently struggle to accurately represent the diversity of human perspectives, often overgeneralizing in favor of majority opinions found in training data. This erasure of minority viewpoints can lead to harmful stereotyping and the disenfranchisement of marginalized groups in downstream applications, such as candidate screening or safety monitoring. While traditional Natural Language Processing often dismisses label ambiguity as "noise," recent scholars suggest these disagreements are valuable signals for fair model development. To address this, my research aims to amplify the variety of viewpoints within LLMs to more accurately reflect genuine human diversity. I investigate how Reinforcement Learning can introduce variation in LLMs via its reward mechanism, beyond objective, verifiable tasks like mathematics or coding.
I hypothesize that a dual-reward mechanism (a reasoning reward to incentivize the exploration of multiple rationales and a measurement of similarity between model outputs and human-annotated label variations) will enable us to capture inherent human diversity across three tasks from social Natural Language Processing. The ultimate goal is to enable robust language technology that transcends majority demographics to reflect the full breadth of human understanding.
Publications by Sebastian Loftus
-
-
Article de journalLoftus, S., Mülthaler, A., Hoeken, S., Zarrieß, S., & Alaçam, Ö. (2025). Using LLMs and Preference Optimization for Agreement-Aware HateWiC Classification. Proceedings of the The 9th Workshop on Online Abuse and Harms (WOAH), 9, 538-547. (Original work published 2025)
-
-
-
Article de journalPeng, S., Sun, Z., Loftus, S., & Plank, B. (2024). Different Tastes of Entities: Investigating Human Label Variation in Named Entity Annotations. Workshop on Understanding Implicit and Underspecified Language, 3, 73-81. https://doi.org/10.18653/v1/2024.unimplicit-1.7 (Original work published 2024)
-