UPDF AI

Does it Make Sense? And Why? A Pilot Study for Sense Making and Explanation

Cunxiang Wang,Shuailong Liang,2 Authors,Tian Gao

2019 · DOI: 10.18653/v1/P19-1393
Annual Meeting of the Association for Computational Linguistics · 112 Citations

TLDR

A benchmark to directly test whether a system can differentiate natural language statements that make sense from those that do not make sense is released and models trained over large-scale language modeling tasks as well as human performance are evaluated, showing that there are different challenges for system sense-making.

Abstract

Introducing common sense to natural language understanding systems has received increasing research attention. It remains a fundamental question on how to evaluate whether a system has the sense-making capability. Existing benchmarks measure common sense knowledge indirectly or without reasoning. In this paper, we release a benchmark to directly test whether a system can differentiate natural language statements that make sense from those that do not make sense. In addition, a system is asked to identify the most crucial reason why a statement does not make sense. We evaluate models trained over large-scale language modeling tasks as well as human performance, showing that there are different challenges for system sense-making.