CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sapienzanlp /dromedario-3-sft-dataset 🐪 Dataset Card for Dromedario 3 📋 Dataset Summary Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.texttext-generation100K<n<1M9 likes508 downloads9d agoHugging Face02sapinsapin /halo-hil halo-hil Web text in hil, re-filtered by language and prepared for pretraining. What changed, and why it had to The earlier version of this dataset was labelled hil by the crawler's own language detection, and that label was never verified. An audit on 2026-09-22 found that most of it was not hil: over a random sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 % Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.tabulartext-generationn<1K0 likes367 downloads2d agoHugging Face03sapienzanlp-course-materials /hw-mnlp-2026 Dataset for Multilingual Natural Language Processing (MNLP) Homeworks This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course. Homework 1 - Semantic Search In the first homework, you are asked to build semantic search systems. You must only use the following variables: query: A single question in natural language. query_id: The question (query) identifier. candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.tabularsentence-similarity10K<n<100K0 likes66 downloads6mo agoHugging Face04sapinsapin /BantayWika BantayWika A FineWeb-compatible pretraining text corpus for Philippine languages, derived from the Bantay-Wika corpus collected by the University of the Philippines Sentro ng Wikang Filipino (UP-SWF) and the UP Digital Signal Processing (DSP) Laboratory. The Bantay-Wika (Language Watch) project was started in 1994 by UP-SWF to track how the Philippine national language is used and develops, particularly in Philippine media. The first phase (1994–2004) involved manual collection and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/BantayWika.texttext-generation10K<n<100K0 likes46 downloads7mo agoHugging Face05tejeshbhalla /sa-pipeline-merged-v2gated sa-pipeline merged training messages Streaming union of five sources, all in a single chat-format messages schema. Designed to give a single SFT corpus that mixes: final-answer-only reasoning (apr30 / v2), explicit <think>...</think> reasoning (opus47, prior_reasoning), agentic tool-calling under the Qwen 3.6 chat template (qwen36_tool_traj), so the trained model learns both to reason and to deliver clean final answers without overfitting to either pattern. Sources… See the full description on the dataset page: https://huggingface.co/datasets/tejeshbhalla/sa-pipeline-merged-v2.texttext-generation100K<n<1M0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.