datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.indaqa
INDAQA - Italian Narrative Dataset for Long-document Question-Answering
INDAQA is the first Italian question-answering dataset specifically designed for long-context Italian narrative texts.
The dataset contains 362 documents paired with reading comprehension questions and reference answers based on Italian literary works sourced from Wikisource.
Questions and answers were automatically generated using Gemini and subsequently underwent both automatic filtering and… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/indaqa.
