Computer Science > Computation and Language

arXiv:2304.07957 (cs)

[Submitted on 17 Apr 2023]

Title:A Question-Answering Approach to Key Value Pair Extraction from Form-like Document Images

Authors:Kai Hu, Zhuoyuan Wu, Zhuoyao Zhong, Weihong Lin, Lei Sun, Qiang Huo

View PDF

Abstract:In this paper, we present a new question-answering (QA) based key-value pair extraction approach, called KVPFormer, to robustly extracting key-value relationships between entities from form-like document images. Specifically, KVPFormer first identifies key entities from all entities in an image with a Transformer encoder, then takes these key entities as \textbf{questions} and feeds them into a Transformer decoder to predict their corresponding \textbf{answers} (i.e., value entities) in parallel. To achieve higher answer prediction accuracy, we propose a coarse-to-fine answer prediction approach further, which first extracts multiple answer candidates for each identified question in the coarse stage and then selects the most likely one among these candidates in the fine stage. In this way, the learning difficulty of answer prediction can be effectively reduced so that the prediction accuracy can be improved. Moreover, we introduce a spatial compatibility attention bias into the self-attention/cross-attention mechanism for \Ours{} to better model the spatial interactions between entities. With these new techniques, our proposed \Ours{} achieves state-of-the-art results on FUNSD and XFUND datasets, outperforming the previous best-performing method by 7.2\% and 13.2\% in F1 score, respectively.

Comments:	AAAI 2023
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2304.07957 [cs.CL]
	(or arXiv:2304.07957v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2304.07957

Submission history

From: Kai Hu [view email]
[v1] Mon, 17 Apr 2023 02:55:31 UTC (1,193 KB)

Computer Science > Computation and Language

Title:A Question-Answering Approach to Key Value Pair Extraction from Form-like Document Images

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:A Question-Answering Approach to Key Value Pair Extraction from Form-like Document Images

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators