Repository logo
Research Data
Publications
Projects
Persons
Organizations
English
Français
Log In(current)
  1. Home
  2. Publications
  3. Article de recherche (journal article)
  4. Stemming Approaches for East European Languages

Stemming Approaches for East European Languages

Author(s)
Dolamic, Ljiljana
Savoy, Jacques  
Institut d'informatique  
Date issued
2008
In
Lecture Notes in Computer Science (LNCS), Springer, 2008/5152//37-44
Abstract
During this CLEF evaluation campaign, the first objective is to propose and evaluate various indexing and search strategies for the Czech language that will hopefully result in more effective retrieval than language-independent approaches (<i>n</i>-gram). Based on the stemming strategy we developed for other languages, we propose that for the Slavic language a light stemmer (inflectional only) and also a second one based on a more aggressive suffix-stripping scheme that will remove some derivational suffixes. Our second objective is to undertake further study of the relative merit of various search engines when exploring Hungarian and Bulgarian documents. To evaluate these solutions we use various effective IR models. Our experiments generally show that for the Bulgarian language, removing certain frequently used derivational suffixes may improve mean average precision. For the Hungarian corpus, applying an automatic decompounding procedure improves the MAP. For the Czech language a comparison of a light and a more aggressive stemmer to remove both inflectional and some derivational suffixes, reveals only small performance differences. For this language only, performance differences between a word-based or a 4-gram indexing strategy are also rather small.
Publication type
journal article
Identifiers
https://libra.unine.ch/handle/20.500.14713/58731
DOI
10.1007/978-3-540-85760-0_4
-
https://libra.unine.ch/handle/123456789/14384
File(s)
Loading...
Thumbnail Image
Download
Name

Dolamic_Ljilana_-_Stemming_Approaches_for_East_European_Languages_20100211.pdf

Type

Main Article

Size

464.51 KB

Format

Adobe PDF

Checksum

(MD5):e6455d9b1b32f86745e05c6f3cc609ff

Université de Neuchâtel logo

Service information scientifique & bibliothèques

Rue Emile-Argand 11

2000 Neuchâtel

contact.libra@unine.ch

Service informatique et télématique

Rue Emile-Argand 11

Bâtiment B, rez-de-chaussée

Powered by DSpace-CRIS

v2.0.0

© 2025 Université de Neuchâtel

Portal overviewUser guideOpen Access strategyOpen Access directive Research at UniNE Open Access ORCIDWhat's new