Logo du site
  • English
  • Français
  • Se connecter
Logo du site
  • English
  • Français
  • Se connecter
  1. Accueil
  2. Université de Neuchâtel
  3. Publications
  4. Stemming Approaches for East European Languages
 
  • Details
Options
Vignette d'image

Stemming Approaches for East European Languages

Auteur(s)
Dolamic, Ljiljana
Editeur(s)
Savoy, Jacques 
Institut d'informatique 
Date de parution
2008
In
Lecture Notes in Computer Science (LNCS), Springer, 2008/5152//37-44
Résumé
During this CLEF evaluation campaign, the first objective is to propose and evaluate various indexing and search strategies for the Czech language that will hopefully result in more effective retrieval than language-independent approaches (<i>n</i>-gram). Based on the stemming strategy we developed for other languages, we propose that for the Slavic language a light stemmer (inflectional only) and also a second one based on a more aggressive suffix-stripping scheme that will remove some derivational suffixes. Our second objective is to undertake further study of the relative merit of various search engines when exploring Hungarian and Bulgarian documents. To evaluate these solutions we use various effective IR models. Our experiments generally show that for the Bulgarian language, removing certain frequently used derivational suffixes may improve mean average precision. For the Hungarian corpus, applying an automatic decompounding procedure improves the MAP. For the Czech language a comparison of a light and a more aggressive stemmer to remove both inflectional and some derivational suffixes, reveals only small performance differences. For this language only, performance differences between a word-based or a 4-gram indexing strategy are also rather small.
URI
https://libra.unine.ch/handle/123456789/14384
DOI
10.1007/978-3-540-85760-0_4
Autre version
http://dx.doi.org/10.1007/978-3-540-85760-0_4
Type de publication
Resource Types::text::journal::journal article
Dossier(s) à télécharger
 main article: Dolamic_Ljilana_-_Stemming_Approaches_for_East_European_Languages_20100211.pdf (464.51 KB)
google-scholar
Présentation du portailGuide d'utilisationStratégie Open AccessDirective Open Access La recherche à l'UniNE Open Access ORCID

Adresse:
UniNE, Service information scientifique & bibliothèques
Rue Emile-Argand 11
2000 Neuchâtel

Construit avec Logiciel DSpace-CRIS Maintenu et optimiser par 4Sciences

  • Paramètres des témoins de connexion
  • Politique de protection de la vie privée
  • Licence de l'utilisateur final