TY - GEN
T1 - SemaFor
T2 - 21st ACM International Conference on Information and Knowledge Management, CIKM 2012
AU - Tsatsaronis, George
AU - Varlamis, Iraklis
AU - Nørvåg, Kjetil
PY - 2012
Y1 - 2012
N2 - Traditional document indexing techniques store documents using easily accessible representations, such as inverted indices, which can efficiently scale for large document sets. These structures offer scalable and efficient solutions in text document management tasks, though, they omit the cornerstone of the documents' purpose: meaning. They also neglect semantic relations that bind terms into coherent fragments of text that convey messages. When semantic representations are employed, the documents are mapped to the space of concepts and the similarity measures are adapted appropriately to better fit the retrieval tasks. However, these methods can be slow both at indexing and retrieval time. In this paper we propose SemaFor, an indexing algorithm for text documents, which uses semantic spanning forests constructed from lexical resources, like Wikipedia, and WordNet, and spectral graph theory in order to represent documents for further processing.
AB - Traditional document indexing techniques store documents using easily accessible representations, such as inverted indices, which can efficiently scale for large document sets. These structures offer scalable and efficient solutions in text document management tasks, though, they omit the cornerstone of the documents' purpose: meaning. They also neglect semantic relations that bind terms into coherent fragments of text that convey messages. When semantic representations are employed, the documents are mapped to the space of concepts and the similarity measures are adapted appropriately to better fit the retrieval tasks. However, these methods can be slow both at indexing and retrieval time. In this paper we propose SemaFor, an indexing algorithm for text documents, which uses semantic spanning forests constructed from lexical resources, like Wikipedia, and WordNet, and spectral graph theory in order to represent documents for further processing.
KW - document indexing
KW - semantic graphs
KW - text representation
UR - https://www.scopus.com/pages/publications/84871071778
U2 - 10.1145/2396761.2398499
DO - 10.1145/2396761.2398499
M3 - Contribución a la conferencia
AN - SCOPUS:84871071778
SN - 9781450311564
T3 - ACM International Conference Proceeding Series
SP - 1692
EP - 1696
BT - CIKM 2012 - Proceedings of the 21st ACM International Conference on Information and Knowledge Management
Y2 - 29 October 2012 through 2 November 2012
ER -