Skip to main navigation Skip to search Skip to main content

HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content

  • Pontifícia Universidade Católica do Paraná

Research output: Contribution to Book/Report typesContribution to conference proceedingspeer-review

Fingerprint

Dive into the research topics of 'HARM: Learning Hate-Aware Reward Model for Evaluating Natural Language Explanations of Offensive Content'. Together they form a unique fingerprint.
Sort by

Keyphrases

Computer Science