Passer à la navigation principale Passer à la recherche Passer au contenu principal

HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content

  • Pontifícia Universidade Católica do Paraná

Résultats de recherche: Chapitre dans un livre, rapport, actes de conférenceParticipation à un ouvrage collectif lié à un colloque ou une conférenceRevue par des pairs

Résumé

Explaining why content is hateful using natural language is crucial for fostering transparency in automated content moderation systems. However, evaluating the quality of such explanations remains an open challenge. General-purpose reward models (RMs), commonly used for scoring natural language outputs, are typically optimized for broad notions of safety. We argue that this optimization penalizes situations where references to stereotypes or offensive content are essential for explanations with higher explanatory fidelity. To address this gap, we introduce SBIC-Explain, a human-validated dataset of 370,788 LLM generated NLEs for offensive content, spanning three levels of human-annotated contextual richness: Tier 1: text-only, Tier 2: + classification-aware, and Tier 3: + semantics-informed. We hypothesize that as human-annotated context increases, explanations should lead to higher perceived explanations with higher explanatory fidelity. Yet, we find that existing RMs systematically assign lower scores to more contextually rich (and often more offensive) explanations, revealing a misalignment between model preferences and explanatory fidelity for this context. We propose HARM (Hate-Aware Reward Model), a RM that integrates interpretable signals to better align reward scores with the needs of hate speech explanation. HARM outperforms general-purpose baselines, improving NLE pair-wise preference. Available at: https://github.com/Lorenzo815/HARM.

langue originaleAnglais
titre19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
EditeurAssociation for Computational Linguistics (ACL)
Pages4393-4431
Nombre de pages39
ISBN (Electronique)9798891763869
Les DOIs
étatPublié - 2026
Evénement19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 - Rabat, Maroc
Durée: 24 mars 202629 mars 2026

Série de publications

Nom19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026

Conférence

Conférence19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
Pays/TerritoireMaroc
La villeRabat
période24/03/2629/03/26

Empreinte digitale

Voici les principaux termes ou expressions associés à « HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content ». Ces libellés thématiques sont générés à partir du titre et du résumé de la publication. Ensemble, ils forment une empreinte digitale unique.

Citer cette ressorce