datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dromedario-3-sft-dataset
🐪 Dataset Card for Dromedario 3
📋 Dataset Summary
Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.halo-hil
halo-hil
Web text in hil, re-filtered by language and prepared for
pretraining.
What changed, and why it had to
The earlier version of this dataset was labelled hil by the
crawler's own language detection, and that label was never verified. An audit
on 2026-09-22 found that most of it was not hil: over a random
sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 %
Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog
news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.BantayWika
BantayWika
A FineWeb-compatible pretraining text corpus for Philippine languages, derived from the Bantay-Wika corpus collected by the University of the Philippines Sentro ng Wikang Filipino (UP-SWF) and the UP Digital Signal Processing (DSP) Laboratory.
The Bantay-Wika (Language Watch) project was started in 1994 by UP-SWF to track how the Philippine national language is used and develops, particularly in Philippine media. The first phase (1994–2004) involved manual collection and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/BantayWika.sa-pipeline-merged-v2
sa-pipeline merged training messages
Streaming union of five sources, all in a single chat-format messages schema.
Designed to give a single SFT corpus that mixes:
final-answer-only reasoning (apr30 / v2),
explicit <think>...</think> reasoning (opus47, prior_reasoning),
agentic tool-calling under the Qwen 3.6 chat template (qwen36_tool_traj),
so the trained model learns both to reason and to deliver clean final answers
without overfitting to either pattern.
Sources… See the full description on the dataset page: https://huggingface.co/datasets/tejeshbhalla/sa-pipeline-merged-v2.
