Just like a woman: A comparative analysis of an LLM and human labels on sexism detection

Metztli Ramírez-González, Delia Irazú Hernández-Farías, Manuel Montes-y-Gómez

Resumen


This work presents a comparative analysis of labeling in sexism detection using ambiguous Spanish-language data selected with the Think Twice method. A subset of examples with content related to sexism was annotated by Mexican women from diverse sociocultural backgrounds and contrasted with labels produced by an LLM. The results show low agreement both among humans and between humans and the LLM, reflecting the interpretative variability inherent in subjective tasks. Despite this variability, the model tends to approximate the average human judgment. These findings highlight the need for annotation schemes and classification approaches that account for cultural and linguistic diversity rather than forcing a single correct interpretation in sensitive tasks such as sexism detection.

Texto completo:

PDF