Computer Science > Computation and Language

arXiv:2309.09092 (cs)

[Submitted on 16 Sep 2023]

Title:The Impact of Debiasing on the Performance of Language Models in Downstream Tasks is Underestimated

Authors:Masahiro Kaneko, Danushka Bollegala, Naoaki Okazaki

View PDF

Abstract:Pre-trained language models trained on large-scale data have learned serious levels of social biases. Consequently, various methods have been proposed to debias pre-trained models. Debiasing methods need to mitigate only discriminatory bias information from the pre-trained models, while retaining information that is useful for the downstream tasks. In previous research, whether useful information is retained has been confirmed by the performance of downstream tasks in debiased pre-trained models. On the other hand, it is not clear whether these benchmarks consist of data pertaining to social biases and are appropriate for investigating the impact of debiasing. For example in gender-related social biases, data containing female words (e.g. ``she, female, woman''), male words (e.g. ``he, male, man''), and stereotypical words (e.g. ``nurse, doctor, professor'') are considered to be the most affected by debiasing. If there is not much data containing these words in a benchmark dataset for a target task, there is the possibility of erroneously evaluating the effects of debiasing. In this study, we compare the impact of debiasing on performance across multiple downstream tasks using a wide-range of benchmark datasets that containing female, male, and stereotypical words. Experiments show that the effects of debiasing are consistently \emph{underestimated} across all tasks. Moreover, the effects of debiasing could be reliably evaluated by separately considering instances containing female, male, and stereotypical words than all of the instances in a benchmark dataset.

Comments:	IJCNLP-AACL 2023
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2309.09092 [cs.CL]
	(or arXiv:2309.09092v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2309.09092

Submission history

From: Masahiro Kaneko [view email]
[v1] Sat, 16 Sep 2023 20:25:34 UTC (7,721 KB)

Computer Science > Computation and Language

Title:The Impact of Debiasing on the Performance of Language Models in Downstream Tasks is Underestimated

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:The Impact of Debiasing on the Performance of Language Models in Downstream Tasks is Underestimated

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators