datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.dromedario-3-sft-dataset
🐪 Dataset Card for Dromedario 3
📋 Dataset Summary
Dromedario 3 is a large-scale Italian instruction-tuning dataset derived from the English Tülu 3 SFT mixture through a principled translation pipeline: instructions and responses are classified according to the Natural Instructions taxonomy, classes are manually reviewed for translation safety, and items in validated classes are machine-translated into Italian. The full procedure is described in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/dromedario-3-sft-dataset.halo-hil
halo-hil
Web text in hil, re-filtered by language and prepared for
pretraining.
What changed, and why it had to
The earlier version of this dataset was labelled hil by the
crawler's own language detection, and that label was never verified. An audit
on 2026-09-22 found that most of it was not hil: over a random
sample of 1,499 sentences, GlotLID v3 called 44 % English, 22 %
Filipino/Tagalog and only 12 % Hiligaynon — much of the corpus was Tagalog
news copy and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/halo-hil.mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.arc_italian
ARC - Italian (IT)
This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly.
Dataset Details
The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.boolq_italian
BoolQ - Italian (IT)
This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine.
Dataset Details
The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question.
The dataset includes the following splits:
Train: 9,427 rows
Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.gsm8k_italian
GSM8K - Italian (IT)
This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education.
Dataset Details
The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.ea-mt-benchmark
Dataset Card for EA-MT
EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate.
Here is an example of a simple sentence with a challenging entity mention:
English: "What is the plot of The Catcher in the Rye?"
Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.SA-Prot-annot
SA-Prot-Annot Dataset (Sci-Align)
🌌 The Sciverse Data Foundation
Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research.
Sciverse consists of three core data… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SA-Prot-annot.SAP-9k
SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use
paper: SAP: State-Guided Data Synthesis with Argument Provenance for Multi-Turn Tool Use
Dataset Overview
This is a high-quality, executor-validated multi-turn tool-use dataset designed for training agentic language models on long-horizon function calling tasks. The dataset focuses on argument-level cross-turn dependency grounding, ensuring tool arguments are sourced from verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Zichen1024/SAP-9k.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.hw-mnlp-2026
Dataset for Multilingual Natural Language Processing (MNLP) Homeworks
This dataset serves for both Homework 1 and Homework 2 of the Multilingual Natural Language Processing (MNLP) course.
Homework 1 - Semantic Search
In the first homework, you are asked to build semantic search systems. You must only use the following variables:
query: A single question in natural language.
query_id: The question (query) identifier.
candidate_chunks: List of candidate answers (only one… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp-course-materials/hw-mnlp-2026.big-pickle-409xTrace of Big Pickle, stealth model from OpenCode Zen. (It's been rumored that it's GLM 4.6)
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
SAP-Hypo5
Dataset Card for SAP-Hypo5
SAP-Hypo5 is an open benchmark for LLM-Agent ASR hypothesis correction on dysarthric speech.
Dataset Description
Following HyPoradise, each selected SAP utterance is paired with its reference transcript and the top-5 ASR hypotheses from Whisper-large-v2 fine-tuned on SAP (PD-only challenge release).
Dataset Sources
Repository: https://github.com/xiuwenz2/SAP-Hypo5
Paper: Towards Robust Dysarthric Speech Recognition: LLM-Agent… See the full description on the dataset page: https://huggingface.co/datasets/xiuwenz2/SAP-Hypo5.piqa_italian
PIQA - Italian (IT)
This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world.
Dataset Details
The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.truthful_qa_italian
TruthfulQA - Italian (IT)
This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions.
Dataset Details
The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.winogrande_italian
Winogrande - Italian (IT)
This dataset is an Italian translation of Winogrande. Winogrande is a large-scale dataset for coreference resolution, commonsense reasoning, and world knowledge. It is based on the original Winograd Schema Challenge dataset.
Dataset Details
The dataset consists of almost 40K examples, each containing a sentence with a blank and two possible fill-in-the-blank options. The task is to choose the correct option that correctly fills in the blank based… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/winogrande_italian.yandexq-qa-chatmlChatML formatted version of its5Q/yandex-q.
sciq_italian
SciQ - Italian (IT)
This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge.
Dataset Details
The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.BantayWika
BantayWika
A FineWeb-compatible pretraining text corpus for Philippine languages, derived from the Bantay-Wika corpus collected by the University of the Philippines Sentro ng Wikang Filipino (UP-SWF) and the UP Digital Signal Processing (DSP) Laboratory.
The Bantay-Wika (Language Watch) project was started in 1994 by UP-SWF to track how the Philippine national language is used and develops, particularly in Philippine media. The first phase (1994–2004) involved manual collection and… See the full description on the dataset page: https://huggingface.co/datasets/sapinsapin/BantayWika.simple_bench
📊 Simple Bench Dataset
A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models
Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.diaforge-utc-r-0725
DiaFORGE UTC: Unified Tool-Calling Conversations Dataset
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
The dataset is generated with the data generation engine described in Figure 1. The engine simulates a user agent and an assistant agent in a dialogue, where the user agent has a persona and the… See the full description on the dataset page: https://huggingface.co/datasets/SAP/diaforge-utc-r-0725.grok-4.1-fast-instruct-308xTrace of Grok 4.1 Fast LLM.
WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models.
Data count (Total: 308):
English - 198
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page
gemma-4-31b-it-304xTrace of Gemma 4 31B LLM.
Data count (Total: 304):
English - 194
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
Bhagavad-Gita_Dataset
Srimad Bhagavad Gita Dataset
A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks.
Dataset Details
Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi
Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/Saptak123/Bhagavad-Gita_Dataset.deepseek-v4-flash-instruct-308xTrace of DeepSeek V4 Flash LLM.
WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models.
Data count (Total: 308):
English - 198
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page
gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me).
Data count (Total: 425):
English - 209
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope).
Brought to you by sapbot from Romarchive
grok-4.1-fast-instruct-308x-cot
Grok 4.1 Fast Instruct 308x traces with added russian CoT
Traces generated using RU-CoT-Generator and google/gemma-3-4b-it as CoT generator.
coderppl
CoderPPL
A curated code perplexity evaluation corpus — 9,324 lines of real-world, working code across 24 files and 5 programming languages.
Source: github.com/sapbotgit/code-doodles
This dataset is designed to measure code perplexity (PPL) — how well a language model predicts actual hand-written code across multiple languages and programming paradigms.
Contents
Language
Files
Examples
Python
5
LLM trainers, proxy scanner, fine-tuning tools
JavaScript… See the full description on the dataset page: https://huggingface.co/datasets/sapbot/coderppl.yandexq-qa-100
YandexQ QA (100 subset)
Traces generated using RU-CoT-Generator and liquid/lfm-2.5-8b-a1b as CoT generator.
Dataset based on sapbot/yandexq-qa-chatml.
