datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu_italian
MMLU - Italian (IT)
This dataset is an Italian translation of Massive Multitask Language Understanding (MMLU). MMLU is a dataset that is composed of multiple-choice questions from 57 different topics, including math, science, and social studies. The dataset is designed to evaluate the ability of models to answer questions across a wide range of topics.
Dataset Details
The dataset consists of multiple-choice questions from 57 different topics. Each question is associated… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/mmlu_italian.arc_italian
ARC - Italian (IT)
This dataset is an Italian translation of the AI2 Reasoning Challenge (ARC). ARC is a question-answering dataset that requires an understanding of natural language text and reasoning capabilities to answer questions correctly.
Dataset Details
The dataset consists of multiple-choice questions, where each question is associated with a set of answer choices (up to 5 choices). The task is to choose the correct answer choice based on the context provided in… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/arc_italian.boolq_italian
BoolQ - Italian (IT)
This dataset is an Italian translation of BoolQ. BoolQ is a question-answering dataset composed of user queries issued to a search engine.
Dataset Details
The task is to predict whether the answer to the question is true or false based on the context provided in the question. A text snippet from Wikipedia is provided as the context for each question.
The dataset includes the following splits:
Train: 9,427 rows
Validation: 3,270 rows… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/boolq_italian.gsm8k_italian
GSM8K - Italian (IT)
This dataset is an Italian translation of GSM8K. GSM8K stands for Grade School Math 8K, a dataset for math word problems, which should be easy to solve for people with an elementary school education.
Dataset Details
The dataset consists of math word problems, where each problem is associated with a possible explanation of how to solve it. The task is to generate the answer to the math problem. The dataset is split into a training set and a test set.… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/gsm8k_italian.ea-mt-benchmark
Dataset Card for EA-MT
EA-MT (Entity-Aware Machine Translation) is a multilingual benchmark for evaluating the capabilities of Large Language Models (LLMs) and Machine Translation (MT) models in translating simple sentences with potentially challenging entity mentions, e.g., entities for which a word-for-word translation may not be accurate.
Here is an example of a simple sentence with a challenging entity mention:
English: "What is the plot of The Catcher in the Rye?"
Italian:… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/ea-mt-benchmark.hellaswag_italian
HellaSwag - Italian (IT)
This dataset is an Italian translation of HellaSwag. HellaSwag is a large-scale commonsense reasoning dataset, which requires reading comprehension and commonsense reasoning to predict the correct ending of a sentence.
Dataset Details
The dataset consists of instances containing a context and a multiple-choice question with four possible answers. The task is to predict the correct ending of the sentence. The dataset is split into a training set… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/hellaswag_italian.big-pickle-409xTrace of Big Pickle, stealth model from OpenCode Zen. (It's been rumored that it's GLM 4.6)
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
SAP-Hypo5
Dataset Card for SAP-Hypo5
SAP-Hypo5 is an open benchmark for LLM-Agent ASR hypothesis correction on dysarthric speech.
Dataset Description
Following HyPoradise, each selected SAP utterance is paired with its reference transcript and the top-5 ASR hypotheses from Whisper-large-v2 fine-tuned on SAP (PD-only challenge release).
Dataset Sources
Repository: https://github.com/xiuwenz2/SAP-Hypo5
Paper: Towards Robust Dysarthric Speech Recognition: LLM-Agent… See the full description on the dataset page: https://huggingface.co/datasets/xiuwenz2/SAP-Hypo5.piqa_italian
PIQA - Italian (IT)
This dataset is an Italian translation of PIQA. PIQA stands for Physical Interaction Question Answering, a dataset of questions about common scenarios that require an understanding of the physical world.
Dataset Details
The dataset consists of questions about common scenarios that require an understanding of the physical world. Each question is associated with a correct answer and a distractor. The task is to predict the correct answer to the… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/piqa_italian.truthful_qa_italian
TruthfulQA - Italian (IT)
This dataset is an Italian translation of TruthfulQA. TruthfulQA is a dataset for fact-based question answering, which contains questions that require factual knowledge to answer correctly. These questions are designed so that some humans would answer them incorrectly because of common misconceptions.
Dataset Details
The dataset is a question answering dataset that contains questions that require factual knowledge to answer correctly and avoid… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/truthful_qa_italian.winogrande_italian
Winogrande - Italian (IT)
This dataset is an Italian translation of Winogrande. Winogrande is a large-scale dataset for coreference resolution, commonsense reasoning, and world knowledge. It is based on the original Winograd Schema Challenge dataset.
Dataset Details
The dataset consists of almost 40K examples, each containing a sentence with a blank and two possible fill-in-the-blank options. The task is to choose the correct option that correctly fills in the blank based… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/winogrande_italian.yandexq-qa-chatmlChatML formatted version of its5Q/yandex-q.
sciq_italian
SciQ - Italian (IT)
This dataset is an Italian translation of SciQ. SciQ is a dataset for scientific questions, which were semi-automatically generated from an existing set of questions. The dataset is designed to test the ability of models to answer questions that require scientific knowledge.
Dataset Details
The dataset consists of science-related questions, where each question is associated with a correct answer and three possible distractors. The task is to predict… See the full description on the dataset page: https://huggingface.co/datasets/sapienzanlp/sciq_italian.simple_bench
📊 Simple Bench Dataset
A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models
Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.grok-4.1-fast-instruct-308xTrace of Grok 4.1 Fast LLM.
WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models.
Data count (Total: 308):
English - 198
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page
gemma-4-31b-it-304xTrace of Gemma 4 31B LLM.
Data count (Total: 304):
English - 194
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
deepseek-v4-flash-instruct-308xTrace of DeepSeek V4 Flash LLM.
WARNING: This trace was made WITHOUT reasoning. Use it to finetune only instruct models.
Data count (Total: 308):
English - 198
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page
gemma-3n-4b-distill-smollm2-360m-instruct-425xTrace of Gemma 3n 4B Distill SmolLM2 360M Instruct LLM by sapbot (me).
Data count (Total: 425):
English - 209
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Note: This was added more as a "examples" of this model's outputs. Of course you will not distill a distilled model (I hope).
Brought to you by sapbot from Romarchive
grok-4.1-fast-instruct-308x-cot
Grok 4.1 Fast Instruct 308x traces with added russian CoT
Traces generated using RU-CoT-Generator and google/gemma-3-4b-it as CoT generator.
yandexq-qa-100
YandexQ QA (100 subset)
Traces generated using RU-CoT-Generator and liquid/lfm-2.5-8b-a1b as CoT generator.
Dataset based on sapbot/yandexq-qa-chatml.
breathe-sapo
breath-sapo
A dataset of 250 complex multi-step reasoning problems across 5 domains, each paired with a structured checklist for verifiable answer evaluation.
Dataset Description
Each example contains a detailed multi-step problem and a checklist — a JSON-encoded dict of expected intermediate results with regex patterns for automated verification.
Categories
Category
Description
Engineering Design
Multi-constraint design problems (solar systems, HVAC… See the full description on the dataset page: https://huggingface.co/datasets/yundog/breathe-sapo.llama-3.1-8b-instruct-419xTrace of Llama 3.1 8B Instruct LLM by Meta.
Data count (Total: 419):
English - 204
Russian - 215
Data is presented in ChatML format and each conversation split by newline.
P.S. I do not recommend anyone to use data from this model, because it's.... way too stupid and random at some time, results are not distill-worthy. Go check Gemma 3 12B dataset.
This model was NOT free, and I had to use OpenRouter for it. Crypto donations for future projects like this are available on my personal page
qwen3.5-397b-a17b-218xTrace of Qwen3.5 397B A17B LLM.
Data count (Total: 218):
English - 108
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
yandexq-qa-cot-1k
YandexQ QA (1K subset)
Traces generated using RU-CoT-Generator and google/gemma-3-4b-it as CoT generator.
Dataset based on sapbot/yandexq-qa-chatml.
gpt-oss-20b-500xTrace of gpt-oss 20B LLM made by OpenAI.
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
gemma-3-12b-it-407xTrace of Gemma 3 12B LLM.
Data count (Total: 407):
English - 198
Russian - 209
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
ling-2.6-flash-422xTrace of Ling 2.6 Flash LLM by inclusionAI.
Data count (Total: 422):
English - 206
Russian - 216
Data is presented in ChatML format and each conversation split by newline.
Brought to you by sapbot from Romarchive
ministral-3-14b-439xTrace of Ministral 3 14B LLM by Mistral.
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
lfm2-24b-a2b-427xTrace of LFM2-24B-A2B LLM by LiquidAI.
Data count (Total: 427):
English - 211
Russian - 216
Data is presented in ShareGPT format and each conversation split by newline.
Brought to you by sapbot from Romarchive
aidentist
AIDentist Dataset
A dataset designed for training and fine-tuning language models to answer questions related to the AIDentist dental clinic management system.
This dataset is mainly intended for instruction-tuning (instruction → answer) tasks in Uzbek language.
📌 Dataset Purpose
The goal of this dataset is to help AI models:
Understand questions about the AIDentist platform
Provide accurate answers about system features
Assist users with platform usage
Enable… See the full description on the dataset page: https://huggingface.co/datasets/saparbayev-azizbek/aidentist.
