Abstract
In this paper we present a system, DoLSuD, for the automatic discovery of relevant substructures in a document layout. DoLSuD, Document Layout Substructure Discovery, extracts, analyzes and describes the visual content of structured documents, such as catalogs, in order to discover repeating and distinctive substructures in the document layout and to establish relations between textual and image content. The paper presents the system along with experimental results and the web based service which utilizes the analysis results.
This work was supported by the European Union under the project VIKEF: Virtual Information and Knowledge Environment Framework (http://vikef.net). Image courtesy of IKEA Italia, copyright held by the IKEA group.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Preview
Unable to display preview. Download preview PDF.
Similar content being viewed by others
References
van Ossenbruggen, J., Stamou, G., Pan, J.Z.: Multimedia Annotations and the Semantic Web. In: SWCASE. Proc. of the International Workshop on Semantic Web Case Studies and Best Practices for eBusiness (2005)
Barnard, K., Duygulu, P., Forsyth, D., de Freitas, N., Blei, D., Jordan, M.: Matching words and pictures. Journal of Machine Learning Research 3, 1107–1135 (2002)
Andreatta, C., Lecca, M., Messelodi, S.: Memory-based object recognition in digital images. In: VMV 2005. Proceedings of 10th International Fall Workshop - Vision, Modelling, and Visualization, Erlangen, Germany (November 16-28, 2005)
Bartolini, R., Giovannetti, E., Marchi, S., Montemagni, S., Andreatta, C., Brunelli, R., Stecher, R., Niedere, C.: Ontology learning in multimedia information extraction from product catalogues. In: Staab, S., Svátek, V. (eds.) EKAW 2006. LNCS (LNAI), vol. 4248, Springer, Heidelberg (2006)
Cook, D., Holder, L.: Substructure discovery using minimum description length and background knowledge. Journal of Artificial Intelligence Research 1, 231–255 (1994)
Coble, J., Rathi, R., Cook, D.J., Holder, L.B.: Iterative structure discovery in graph-based data. International Journal on Artificial Intelligence Tools 14(1-2), 101–124 (2005)
Author information
Authors and Affiliations
Editor information
Rights and permissions
Copyright information
© 2007 Springer-Verlag Berlin Heidelberg
About this paper
Cite this paper
Andreatta, C. (2007). Document Layout Substructure Discovery. In: Falcidieno, B., Spagnuolo, M., Avrithis, Y., Kompatsiaris, I., Buitelaar, P. (eds) Semantic Multimedia. SAMT 2007. Lecture Notes in Computer Science, vol 4816. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-77051-0_31
Download citation
DOI: https://doi.org/10.1007/978-3-540-77051-0_31
Publisher Name: Springer, Berlin, Heidelberg
Print ISBN: 978-3-540-77033-6
Online ISBN: 978-3-540-77051-0
eBook Packages: Computer ScienceComputer Science (R0)