Mediaspace scheduled maintenance: Aug 25, 2026 07:00 - 12:00 AM. During this time, videos will be temporarily unavailable. Check status updates.
We propose a novel fully-automated approach towards inducing multilingual taxonomies fromWikipedia. Given an English taxonomy, our approach first leverages the interlanguage links of Wikipedia to automatically construct training datasets for the is-a relation in the target language. Character-level classifiers are trained on the constructed datasets, and used in an optimal path discovery framework to induce high-precision, high-coverage taxonomies in other languages. Through experiments, we demonstrate that our approach significantly outperforms the state-of-the-art, heuristics-heavy approaches for six languages. As a consequence of our work, we release presumably the largest and the most accurate multilingual taxonomic resource spanning over 280 languages.
Simon Nessim Henein, Hubert Pierre-Marie Benoît Schneegans, Florent Cosandier
Tom Ian Battin, Hannes Markus Peter, Tyler Joe Kohler, Susheel Bhanu Busi, Grégoire Marie Octave Edouard Michoud, Stylianos Fodelianakis, Massimo Bourquin, Leïla Ezzat