TY - GEN
T1 - Searching for truth in a database of statistics
AU - Cao, Tien Duc
AU - Manolescu, Ioana
AU - Tannier, Xavier
N1 - Publisher Copyright:
© 2018 Association for Computing Machinery.
PY - 2018/6/10
Y1 - 2018/6/10
N2 - The proliferation of falsehood and misinformation, in particular through the Web, has lead to increasing energy being invested into journalistic fact-checking. Fact-checking journalists typically check the accuracy of a claim against some trusted data source. Statistic databases such as those compiled by state agencies are often used as trusted data sources, as they contain valuable, high-quality information. However, their usability is limited when they are shared in a format such as HTML or spreadsheets: this makes it hard to find the most relevant dataset for checking a specific claim, or to quickly extract from a dataset the best answer to a given query. We present a novel algorithm enabling the exploitation of such statistic tables, by (i) identifying the statistic datasets most relevant for a given fact-checking query, and (ii) extracting from each dataset the best specific (precise) query answer it may contain. We have implemented our approach and experimented on the complete corpus of statistics obtained from INSEE, the French national statistic institute. Our experiments and comparisons demonstrate the effectiveness of our proposed method.
AB - The proliferation of falsehood and misinformation, in particular through the Web, has lead to increasing energy being invested into journalistic fact-checking. Fact-checking journalists typically check the accuracy of a claim against some trusted data source. Statistic databases such as those compiled by state agencies are often used as trusted data sources, as they contain valuable, high-quality information. However, their usability is limited when they are shared in a format such as HTML or spreadsheets: this makes it hard to find the most relevant dataset for checking a specific claim, or to quickly extract from a dataset the best answer to a given query. We present a novel algorithm enabling the exploitation of such statistic tables, by (i) identifying the statistic datasets most relevant for a given fact-checking query, and (ii) extracting from each dataset the best specific (precise) query answer it may contain. We have implemented our approach and experimented on the complete corpus of statistics obtained from INSEE, the French national statistic institute. Our experiments and comparisons demonstrate the effectiveness of our proposed method.
U2 - 10.1145/3201463.3201467
DO - 10.1145/3201463.3201467
M3 - Conference contribution
AN - SCOPUS:85050086580
T3 - Proceedings of the 21st Workshop on the Web and Databases, WebDB 2018
BT - Proceedings of the 21st Workshop on the Web and Databases, WebDB 2018
PB - Association for Computing Machinery
T2 - 21st Workshop on the Web and Databases, WebDB 2018
Y2 - 10 June 2017 through 10 June 2017
ER -