Part of Speech n-Grams and Information Retrieval

Publikation: Bidrag til tidsskriftTidsskriftartikelForskningfagfællebedømt

Efforts to use linguistics in information retrieval (IR) were initiated in the 1980s, and intensified in the 1990s, reporting performance benefits (see the overviews by Smeaton 1986 & 1999, Karlgren 1993, and Tait 2005). After that time, these efforts decreased: baseline system performance improved, and the cost associated with linguistic processing was not worth the small benefits over the already improved baselines (Tait, 2005). At present, most research on linguistics for IR tends to be geared towards domain-specific IR applications that seem to benefit more from linguistics, like question-answering (Tait & Oakes 2006). Although such applications are important, they should not limit the scope of research into linguistics for IR. In this work, we present an alternative use of linguistics, part of speech information in particular, to compute a term weight of informative content. This term weight is a novel application of linguistics to IR, and can benefit retrieval performance of general IR systems.
OriginalsprogEngelsk
TidsskriftRevue Francaise de Linguistique Appliquee
Vol/bindXIII
Udgave nummer2008/1
Sider (fra-til)9-22
ISSN1386-1204
StatusUdgivet - 2008
Eksternt udgivetJa

ID: 38240584