TY - GEN
T1 - A Transformer-based Siamese Network For Word Image Retrieval In Historical Documents
AU - Fathallah, Abir
AU - El-Yacoubi, Mounim A.
AU - Essoukri Ben Amara, Najoua
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023/1/1
Y1 - 2023/1/1
N2 - The increasing availability of digitized historical documents has sparked a need for effective information processing tools to extract the valuable information contained within them. Word spotting, an area of focus in historical document analysis, involves identifying specific words within images of documents. In this paper, we propose a novel approach for word spotting in historical Arabic documents, utilizing improved feature representations for learning word images. More precisely, we put forward an end-to-end approach for generating word image descriptors, based on the Siamese vision transformer architectures. The model learning is guided by a contrastive loss objective. Additionally, we carry out transfer learning techniques by leveraging knowledge acquired from two distinct source domains to generalize model learning. The proposed approach utilizes the embedding space to evolve the word spotting system by projecting the query word image and all reference word images into the embedding space, where their similarity is determined based on their corresponding embedding vectors. Our method is evaluated on the historical Arabic VML-HD dataset and the results indicate that our approach significantly outperforms state-of-the-art methods.
AB - The increasing availability of digitized historical documents has sparked a need for effective information processing tools to extract the valuable information contained within them. Word spotting, an area of focus in historical document analysis, involves identifying specific words within images of documents. In this paper, we propose a novel approach for word spotting in historical Arabic documents, utilizing improved feature representations for learning word images. More precisely, we put forward an end-to-end approach for generating word image descriptors, based on the Siamese vision transformer architectures. The model learning is guided by a contrastive loss objective. Additionally, we carry out transfer learning techniques by leveraging knowledge acquired from two distinct source domains to generalize model learning. The proposed approach utilizes the embedding space to evolve the word spotting system by projecting the query word image and all reference word images into the embedding space, where their similarity is determined based on their corresponding embedding vectors. Our method is evaluated on the historical Arabic VML-HD dataset and the results indicate that our approach significantly outperforms state-of-the-art methods.
KW - Historical Arabic documents
KW - Learning representation
KW - Siamese network
KW - Transfer learning
KW - Vision transformer
KW - Word spotting
U2 - 10.1109/SWC57546.2023.10449180
DO - 10.1109/SWC57546.2023.10449180
M3 - Conference contribution
AN - SCOPUS:85187396998
T3 - Proceedings - 2023 IEEE SmartWorld, Ubiquitous Intelligence and Computing, Autonomous and Trusted Vehicles, Scalable Computing and Communications, Digital Twin, Privacy Computing and Data Security, Metaverse, SmartWorld/UIC/ATC/ScalCom/DigitalTwin/PCDS/Metaverse 2023
BT - Proceedings - 2023 IEEE SmartWorld, Ubiquitous Intelligence and Computing, Autonomous and Trusted Vehicles, Scalable Computing and Communications, Digital Twin, Privacy Computing and Data Security, Metaverse, SmartWorld/UIC/ATC/ScalCom/DigitalTwin/PCDS/Metaverse 2023
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 9th IEEE Smart World Congress, SWC 2023
Y2 - 28 August 2023 through 31 August 2023
ER -