Preaload Image
Back

LET'S BRING YOUR IDEAS TO LIFE

LANGUAGE
PROJECTS

Empowering Language Intelligence via Citizen
Science-Driven Lexicon Development

Home > Research Activities > LiveLanguage: Collaboration Opportunities

LiveLanguage: Vision and Objectives

LiveLanguage is an initiative for the collaborative research and development of computational language resources and tools for underserved languages, on ethical grounds. The aims of LiveLanguage are to provide technical and methodological support for such collaborative efforts and their dissemination. According to the ethics policy of LiveLanguage, local communities set the goals of these collaborative projects and keep the intellectual property of their results. While these principles constitute an ethical minimum for a balanced collaborative relationship, they are also justified by efficiency: they ensure that the project is useful and relevant, and that it motivates the local community in engaging with the project. A local institutional framework greatly simplifies the collaboration process: the local institution can act as the IP owner and has the necessary structure and network of people to organise local efforts. 

Research Goals

The following are past, current, and future areas of research pursued in the context of LiveLanguage. The initiative is not limited, however, to these areas and welcomes researchers with different goals.

Culture-aware Language Resources

By “culture-aware” we mean language resources that focus on what makes languages unique and culturally specific in terms of content that they describe. Rather than treating languages as different forms of expressing meaning that is seen as universal, we consider each language as having a unique perspective on the world. Culture-aware language resources are lexicons, corpora, or other kinds of resources that focus on capturing that unique perspective of each language. These resources then feed into data-driven NLP applications in order to improve their sensitivity and applicability to language- and culture-specific phenomena.

Culture-aware Machine Translation

Even the most sophisticated machine translation tools tend to fail on linguistically or culturally specific, hard-to-translate terms and expressions. Yet, such terms and phrases often express particularly relevant notions of local culture. We aim to evaluate and improve machine translation systems over such expressions.

Games for Learning Endangered Languages

Of the roughly 7000 languages of the world, the vast majority are endangered with decreasing numbers of speakers. In many such communities, there is a genuine interest in transmitting the language and the culture to younger generations but, due to circumstances such as geographic isolation or economic vulnerability, there is a lack of adequate educational material. Our aim is to exploit our culture-aware language resources for the creation of educational games for language learning.

Culture-aware Language Models

While the best multilingual large language models may support hundreds of languages, their performance drops on languages for which training corpora are scarce. There are many efforts nowadays on the development of language models for under-resourced languages. Our approach is to rely on the combination of our culture-aware language resources with knowledge-driven methods (e.g. knowledge graphs) to enhance the contents of language models and thus to boost their abilities on these languages

Resource Development Projects

Principles
When collecting linguistic data and building language resources for a given language spoken by a given community, LiveLanguage imposes two major constraints on the working method.

  • Priority to data quality: the correctness of the data collected must be validated, and only high-quality resources will be built. In practice, this means always putting human validators in the loop (experts or multiple native speakers), even if certain data collection techniques can be automated. Human contribution and oversight are key for data quality.
  • Local control over goals and results: the speaker community (or their representatives, e.g. local researchers) determines the goals based on local needs. At the end of the effort, results will be owned, in terms of intellectual property, by the local community, even if a larger-scale sharing of results is encouraged via appropriately chosen licensing agreements.

Methodology

Based on the principles above, LiveLanguage resource building efforts are aligned with the following high-level methodology:

  1. Project specification based on local needs, driven by a local institution (university, research centre), i.e. You:
    1. goals and motivations for the project;
    2. languages and domains to be covered by the project;
    3. institutions and actors involved;
    4. tools and infrastructure required.
  2. local deployment of diversity-aware supporting tools, supported and (in part) provided by DataScientia;
  3. local resource development, organised and driven by the local institution;
  4. local dissemination and exploitation of results by the local institution;
  5. (optional) sharing of results with LiveLanguage;
  6. (optional) global dissemination and exploitation, publishing results on the LiveLanguage online catalogue and possibly integrating them with other multilingual resources.

Propose a new Project

LiveLanguage resource development efforts are not limited to lexicons: we have ongoing efforts on collecting text corpora as well as culture-aware parallel corpora for machine translation. We are open to any kind of effort that respects the overall principles of the initiative, outlined above.

LiveLanguage Resource Development Projects

Example: Culture-Aware Lexicon Development

Why? Lexical resources—that record the words of a language and their meanings—are widely used by humans as well as by computational applications. Among others, they can be used for translation (human or automated), for language learning, for language understanding (e.g. word sense disambiguation), for cross-lingual studies in linguistics or cognitive science, but also as corpora to feed language models. We are interested in the creation of new lexicons as well as in the validation of existing ones: there are lots of digital lexical resources in circulation that are of low quality, yet people continue using them, especially non-speakers working on multilingual projects. In particular, wordnets can be of very high or very low quality, depending on the methods used for building them: automatically-built wordnets often present major quality issues.

Studying cross-linguistic diversity: cognate clusters for the concept of book, computed from UKC and CogNet data.

What? A lexical resource can focus on the basic lexicon of a language, covering so-called basic-level concepts, or on specialised terminology, i.e. the concepts of a given discipline or domain. A third type of lexical knowledge, of especial relevance to us, are culturally relevant concepts, as well as evidence of lexical untranslatability—when a lexeme is not directly translatable to a lexeme in another language, requiring additional explanation or approximation to convey its meaning. It is the explicit coverage of language-specific and culture-specific terms, as well as of lexical untranslatability, that our lexicon development approach differs from other efforts.

How? LiveLanguage lexicon development efforts follow the six-step methodology outlined above, as shown in the figure.

The LiveLanguage resource development methodology applied to example lexical resources over German-originated languages and dialects of the Italian Alps.

The methodology is articulated around the Universal Knowledge Core (UKC), a culture-aware lexical database covering over 2,000 languages. Lexicon development can be bootstrapped from existing UKC data, and results are typically re-integrated into the UKC in order to provide cross-lingual connectivity (translation, evidence for untranslatability, other types of cross-linguistic knowledge). In terms of data collection methods, projects can possibly use expert-sourcing, crowdsourcing, as well as corpus-based automated methods as long as strong human validation ensures the correctness of its results.

Past and Running Resource Development Projects

  • Unified Scottish Gaelic wordnet
  • Mongolian wordnet
  • Setswana basic lexicon and space domain terms
    • People: Tebatso Moape, Sunday O. Ojo, Gábor Bella, Abed Alhakim Freihat
    • Institutions: Durban University of Technology, University of Trento
    • Publications: n/a
    • Website: n/a
  • Arabic wordnet v3
  • IndoUKC
  • KinDiv: a massively multilingual lexical database on kinship terms
    • People: Temuulen Khishigsuren, Gábor Bella, Khuyagbaatar Batsuren, Abed Alhakim Freihat, Nandu Chandran Nair, Amarsanaa Ganbold, Hadi Khalilia, Yamini Chandrashekar
    • Institutions: National University of Mongolia, University of Trento
    • Publications:
    • Website:
  • Kinship terms in Arabic dialects
  • Basic and kinship terms in Indonesian, Banjar, and Javanese
  • Thai wordnet quality assessment
    • People: Kittiporn Theamnooch, Gábor Bella
    • Institutions: Kasetsart University
    • Publications: n/a
    • Website: n/a
  • Basic level categories in nine languages (Arabic, Turkish, Indonesian, Persian Javanese, Banjarese, Ukrainian, Punjabi, Urdu).
    • People: Hadi Khalilia, Gábor Bella, Rusma Noortyani, Muhammad Irfan, Tetiana Bihun, Shandy Darma, Arya Torabi, Kybra Korkmaz, Sehrish Habib, 
    •  and Fausto Giunchiglia
    • Institutions: University of Trento, Palestine Technical University – Kadoorie (PTUK), IMT Atlantique, Universitas Lambung Mangkurat (ULM), University of Education-Lahore.
    • Publications: n/a
    • Website: n/a

How it Works?

The LiveLanguage initiative is happy to help with your project, by providing tools, expertise, or infrastructure. You are encouraged to get in contact with us. Your project may be similar to one of our past projects in terms of objectives or method, or it may be entirely new. In each potential new collaboration, the first input we are typically looking for are the needs and objectives of the speaker community your project is targeting. Once a first contact is established, in order to make your intentions concrete, we typically ask our partners to describe their intentions via submitting the proposal for the project (available for download below). 

Download collaboration proposal template

DR. Hadi Khalilia

LiveLanguage Coordinator