Skip to content
GuaranIA

Bringing the Guarani into the digital world through AI

GuaranIA

GuaranIA is a project funded by the IDB Lab and developed by CIDIT to advance digital inclusion for Guarani-speaking communities through open AI infrastructure, including models, datasets, tools, and applications that improve access to digital products and services.

GuaranIA has also a strong social focus, with community engagement and participatory research to ensure that the produced tools are culturally appropriate, linguistically accurate, and socially beneficial. Learn more about the fieldwork activities.

Technical Contributions

LLMs, Datasets, and Tools for Guarani.

Project's technical outcomes for building AI infrastructure in low-resource language settings.

Models

Model

GuaranIA-Gemma

In progress

Open-weight LLM based on Google Gemma-4 and adapted for Guarani tasks. To be publicly available on HuggingFace soon.

TBD

Model

GuaranIA-Llama

In progress

Open-weight LLM based on META Llama-3.1 and adapted for Guarani tasks. To be publicly available on HuggingFace soon.

TBD

Datasets

Corpus

Kuatia

Beta

Curated collection of more than 40 real and synthetic sources of Guarani documents, including articles, websites, books, and social media content.

Benchmark

MMLU (Guarani)

Experimental

Professionally translated version of the MMLU Lite benchmark for Guarani. To be publicly available on HuggingFace soon.

CC-BY-4.0

Benchmark

MGSM (Guarani)

In progress

Participatory translated version of the Multilingual Grade School Math (MGSM) benchmark for Guarani. To be publicly available on HuggingFace soon.

TBD

Benchmark

WLNI (Guarani)

In progress

Participatory translated version of the Winograd benchmark for Guarani. To be publicly available on HuggingFace soon.

TBD

Tools & Apps

Package

GuaraScrapper

Beta

Web crawler designed to traverse public websites and collect textual content written in Guarani.

Apache-2.0View on GitHub

Package

GuaraTune

Beta

Framework for adapting open-weight models to Guarani through CPT and SFT on Guarani corpora.

Web Application

Mboehara Digital

Beta

Platform to facilitate the execution of participatory translation and annotation initiatives.

Mobile Application

AIContigo

In progress

AI-powered WhatsApp-based app to help patients follow their medical treatments.

TBD

Methodologies

Methodology

Corpus Pipeline

In progress

Methodology for creating Kuatia, the Guarani corpus, through state-of-the-art pre-processing techniques and language-experts validation.

TBD

Publications

Research, reports, and project writing.

Public articles, papers, release notes, and technical reports produced throughout the project.

Medium article

How do modern AI tools understand Guarani?

A blog article introducing experiments to study how state-of-the-art automatic speech recognition (ASR) tools handle Guarani, comparing the performance of five models, including META’s Omnilingual ASR.

Medium article

How well do today’s AI models handle Guarani?

A blog article introducing a comparative study that evaluates 14 state-of-the-art generative AI models on Guarani translation tasks, revealing a clear trade-off between translation fidelity and linguistic validity.

GuaranIA team

[some of the] people making GuaranIA possible.

Luca Cernuzzi

Luca Cernuzzi

Project Coordinator

Software engineering professor and senior researcher with doctorates from Italian universities, 130+ publications, and extensive R&D, consulting, and public-sector advisory experience.

Jorge Saldivar

Jorge Saldivar

Senior AI Research Engineer

Computer Scientist with a Ph.D. in Information and Communication Technologies from the University of Trento, Italy and experience in AI, NLP, Data Science, and Computational Social Science.

Marvin Aguero

Marvin Aguero

Senior AI Research Engineer

NLP engineering researcher with a PhD in Computer Science from the University of Granada, Spain, and a specialization in text mining, machine learning, and data processing for low-resource languages.

Teresa Gamarra

Teresa Gamarra

Community Coordinator

Senior risk, communication, and social policy expert with 35+ years leading humanitarian, public-sector, research, and international cooperation projects.

Luis Chiruzzo

Luis Chiruzzo

AI/NLP Technical Collaborator

Assistant professor at Universidad de la República (UdelaR), Uruguay, with Computer Science Engineering training and MSc and PhD degrees from UdelaR. Expert in NLP for low-resource languages

Agustin Lucas

Agustin Lucas

AI Research Engineer

Machine Learning Engineer specializing in production GenAI, NLP, RAG, multimodal systems, low-resource translation, and cloud-based ML infrastructure.

Andrea Báez

Andrea Báez

AI Research Engineer

Computer Scientist and Information Engineer with expertise in AI, machine learning and data science.

Cecilia González

Cecilia González

Software Developer

Computer engineering student and Python developer experienced in REST APIs and NLP projects.

Carlos Lugo

Carlos Lugo

Software Engineer

Software Engineer with experience in backend and frontend web development, and machine learning.

Raquel Insfrán

Raquel Insfrán

Software Developer

Computer Engineering student and certified MERN developer skilled in web applications, databases, teamwork, creative problem-solving, and continuous learning.

Ramón Araujo

Ramón Araujo

Software Developer

Computer Engineering student with expertise in software development, data processing and analysis.

Founders and partners

IDB Lab logo
Microsoft logo
Universidad de la República logo
Secretaría de Políticas Lingüísticas logo
Ministerio de Educación y Ciencia logo
UNI logo
Ycuá Bolaños logo
Ayacape logo
Cruz Roja Paraguay logo