BECEL: Benchmark for Consistency Evaluation of Language Models

Jang, Myeongjun; Kwon, Deuk Sin; Lukasiewicz, Thomas

Record link:

http://hdl.handle.net/20.500.12708/192675

Title:

BECEL: Benchmark for Consistency Evaluation of Language Models

Citation:

Jang, M., Kwon, D. S., & Lukasiewicz, T. (2022). BECEL: Benchmark for Consistency Evaluation of Language Models. In N. Calzolari, C.-R. Huang, & H. Kim (Eds.), Proceedings of the 29th International Conference on Computational Linguistics (pp. 3680–3696). International Committee on Computational Linguistics. http://hdl.handle.net/20.500.12708/192675

Publication Type:

Inproceedings - Full-Paper Contribution

Language:

English

Authors:

Jang, Myeongjun
Kwon, Deuk Sin
Lukasiewicz, Thomas

Organisational Unit:

E192-07 - Forschungsbereich Artificial Intelligence Techniques
E192-03 - Forschungsbereich Knowledge Based Systems

Published in:

Proceedings of the 29th International Conference on Computational Linguistics

Date (published):

Oct-2022

Event name:

29th International Conference on Computational Linguistics

Event date:

12-Oct-2022 - 17-Oct-2022

Event place:

Gyeongju, Korea (the Republic of)

Number of Pages:

Publisher:

International Committee on Computational Linguistics

Keywords:

behavioural consistency; language models

Abstract:

Behavioural consistency is a critical condition for a language model (LM) to become trustworthy like humans. Despite its importance, however, there is little consensus on the definition of LM consistency, resulting in different definitions across many studies. In this paper, we first propose the idea of LM consistency based on behavioural consistency and establish a taxonomy that classifies previously studied consistencies into several sub-categories. Next, we create a new benchmark that allows us to evaluate a model on 19 test cases, distinguished by multiple types of consistency and diverse downstream tasks. Through extensive experiments on the new benchmark, we ascertain that none of the modern pre-trained language models (PLMs) performs well in every test case, while exhibiting high inconsistency in many cases. Our experimental results suggest that a unified benchmark that covers broad aspects (i.e., multiple consistency types and tasks) is essential for a more precise evaluation.

Research Areas:

Information Systems Engineering: 100%

Science Branch:

1020 - Informatik: 80%
1010 - Mathematik: 20%

Appears in Collections:

Conference Paper

Show full item record

Page view(s)

217

checked on Jan 25, 2024

Google Scholar^TM

Check

Page view(s)

Google ScholarTM

Google Scholar^TM