当前位置: X-MOL 学术J. Web Semant. › 论文详情
Our official English website, www.x-mol.net, welcomes your feedback! (Note: you will need to create a separate account there.)
From tabular data to knowledge graphs: A survey of semantic table interpretation tasks and methods
Journal of Web Semantics ( IF 2.5 ) Pub Date : 2022-11-19 , DOI: 10.1016/j.websem.2022.100761
Jixiong Liu , Yoan Chabot , Raphaël Troncy , Viet-Phi Huynh , Thomas Labbé , Pierre Monnin

Tabular data often refers to data that is organized in a table with rows and columns. We observe that this data format is widely used on the Web and within enterprise data repositories. Tables potentially contain rich semantic information that still needs to be interpreted. The process of extracting meaningful information out of tabular data with respect to a semantic artefact, such as an ontology or a knowledge graph, is often referred to as Semantic Table Interpretation (STI) or Semantic Table Annotation. In this survey paper, we aim to provide a comprehensive and up-to-date state-of-the-art review of the different tasks and methods that have been proposed so far to perform STI. First, we propose a new categorization that reflects the heterogeneity of table types that one can encounter, revealing different challenges that need to be addressed. Next, we define five major sub-tasks that STI deals with even if the literature has mostly focused on three sub-tasks so far. We review and group the many approaches that have been proposed into three macro families and we discuss their performance and limitations with respect to the various datasets and benchmarks proposed by the community. Finally, we detail what are the remaining scientific barriers to be able to truly automatically interpret any type of tables that can be found in the wild Web.



中文翻译:

从表格数据到知识图:语义表解释任务和方法的调查

表格数据通常是指在具有行和列的表中组织的数据。我们观察到这种数据格式广泛用于 Web 和企业数据存储库中。表可能包含仍然需要解释的丰富语义信息。从关于语义人工制品(例如本体或知识图)的表格数据中提取有意义信息的过程通常称为语义表解释 (STI) 或语义表注释。在这份调查报告中,我们的目标是对迄今为止提出的执行 STI 的不同任务和方法进行全面和最新的最新审查。首先,我们提出了一种新的分类,反映了人们可能遇到的表类型的异质性,揭示了需要解决的不同挑战。接下来,我们定义了 STI 处理的五个主要子任务,即使迄今为止文献主要集中在三个子任务上。我们回顾了已提出的许多方法并将其分为三个宏系列,并讨论了它们相对于社区提出的各种数据集和基准的性能和局限性。最后,我们详细说明了能够真正自动解释可在野生网络中找到的任何类型的表格的剩余科学障碍。

更新日期:2022-11-19
down
wechat
bug