Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards

Arzt, Varvara; Hanbury, Allan

doi:10.18653/v1/2024.genbench-1.8

Record link:

http://hdl.handle.net/20.500.12708/209895

Title:

Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards

Citation:

Arzt, V., & Hanbury, A. (2024). Beyond the Numbers: Transparency in Relation Extraction Benchmark Creation and Leaderboards. In D. Hupkes, V. Dankers, K. Batsuren, A. Kazemnejad, C. Christodoulopoulos, M. Giulianelli, & R. Cotterel (Eds.), Proceedings of the 2nd GenBench Workshop on Generalisation (Benchmarking) in NLP (pp. 120–130). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.genbench-1.8

Publisher DOI:

10.18653/v1/2024.genbench-1.8

Publication Type:

Inproceedings - Full-Paper Contribution

Language:

English

Authors:

Arzt, Varvara
Hanbury, Allan

Organisational Unit:

E194-04 - Forschungsbereich Data Science

Published in:

Proceedings of the 2nd GenBench Workshop on Generalisation (Benchmarking) in NLP

ISBN:

979-8-89176-182-7

DOI of the book:

10.18653/v1/2024.genbench-1

Date (published):

2024

Event name:

GenBench Workshop 2024

Event date:

16-Nov-2024

Event place:

Miami, United States of America (the)

Number of Pages:

Publisher:

Association for Computational Linguistics

Peer reviewed:

Yes

Keywords:

Relation Extraction; Benchmarks; Leaderboards; Transparency

Abstract:

This paper investigates the transparency in the creation of benchmarks and the use of leaderboards for measuring progress in NLP, with a focus on the relation extraction (RE) task. Existing RE benchmarks often suffer from insufficient documentation, lacking crucial details such as data sources, inter-annotator agreement, the algorithms used for the selection of instances for datasets, and information on potential biases like dataset imbalance. Progress in RE is frequently measured by leaderboards that rank systems based on evaluation methods, typically limited to aggregate metrics like F1-score. However, the absence of detailed performance analysis beyond these metrics can obscure the true generalisation capabilities of models. Our analysis reveals that widely used RE benchmarks, such as TACRED and NYT, tend to be highly imbalanced and contain noisy labels. Moreover, the lack of class-based performance metrics fails to accurately reflect model performance across datasets with a large number of relation types. These limitations should be carefully considered when reporting progress in RE. While our discussion centers on the transparency of RE benchmarks and leaderboards, the observations we discuss are broadly applicable to other NLP tasks as well. Rather than undermining the significance and value of existing RE benchmarks and the development of new models, this paper advocates for improved documentation and more rigorous evaluation to advance the field.

Research Areas:

Information Systems Engineering: 100%

Science Branch:

6020 - Sprach- und Literaturwissenschaften: 10%
1020 - Informatik: 90%

Appears in Collections:

Conference Paper

Show full item record

Page view(s)

118

checked on Jan 28, 2025

Google Scholar^TM

Check

Page view(s)

Google ScholarTM

Google Scholar^TM