datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CulturaY
CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages
Dataset Summary
From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset.
Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies.
This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.zebra-cot-mistral-small-3.2-24b-preprocessed
Zebra-CoT Preprocessed — Mistral Hackathon 2026
Preprocessed version of the Zebra-CoT dataset for fine-tuning Mistral-Small-3.2-24B-Instruct.
Format
text: formatted as [INST] question [/INST] <think> reasoning </think> answer
image: PIL JPEG image for the corresponding visual task
Usage
Fine-tuning Mistral-Small-3.2-24B on chain-of-thought visual reasoning.
Hackathon
Created for Mistral Hackaton 2026 — Fine-tuning track with W&B.
Snorkel-Mistral-PairRM-DPO-Dataset
Dataset:
This is the data used for training Snorkel model
We use ONLY the prompts from UltraFeedback; no external LLM responses used.
Methodology:
Generate 5 response variations for each prompt from a subset of 20,000 using the LLM - to start, we used Mistral-7B-Instruct-v0.2.
Apply PairRM for response reranking.
Update the LLM by applying Direct Preference Optimization (DPO) on the top (chosen) and bottom (rejected) responses.
Use this LLM as the base model for the next… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Snorkel-Mistral-PairRM-DPO-Dataset.mistral_playwright_mcp
Playwright MCP Fine-Tuning Dataset for Mistral 7B Instruct 0.3
A testing-focused dataset for fine-tuning Mistral 7B Instruct 0.3 on Playwright MCP (Model Context Protocol) browser automation tasks. This dataset has been filtered to include only crisp, actionable browser testing scenarios.
Dataset Description
This dataset contains 1,235 examples of browser automation testing tasks formatted for instruction-following models. It combines data from multiple sources including… See the full description on the dataset page: https://huggingface.co/datasets/rzagirov/mistral_playwright_mcp.based-chat-v0.1-Mistral-Nemo-Base-2407
Based-Chat v0.1 (Mistral Nemo Base 2407)
This dataset was developed as part of an exploration into understanding the necessity of supervised datasets for fine-tuning base LLMs into conversational models.
It's a synthetic dataset created with Mistral-Nemo-Base-2407, and used to fine-tune that model, producing relay-v0.1-Mistral-Nemo-2407.
Methodology
This synthetic dataset is generated using the following as conversation starters:
facebook/empathetic_dialogues… See the full description on the dataset page: https://huggingface.co/datasets/danlou/based-chat-v0.1-Mistral-Nemo-Base-2407.stanford-encyclopedia-of-philosophy_chat_multi_turn_mistral_largeThis dataset is essentially identical to the Stanford Encyclopedia of Philosophy Chat Multi-turn Dataset, with one key difference: it uses Mistral Large 2 for conversation generation instead of LLaMA 3.1 70B.
All other aspects, including format, statistics, and intended use, remain the same as the original dataset.
robuchan-data
Robuchan Dataset
Synthetic dietary recipe adaptation dataset for fine-tuning language models. Each example is a chat-format conversation where a user provides a recipe and dietary restriction, and the assistant produces a structured adaptation.
Generated for the Mistral AI Worldwide Hackathon Tokyo (Feb 28 - Mar 1, 2026).
Associated model: sumitdotml/robuchan
Dataset Structure
Splits
Split
Rows
Purpose
train
1,090
Fine-tuning training set… See the full description on the dataset page: https://huggingface.co/datasets/mistral-hackaton-2026/robuchan-data.mistral_ultrafeedback_unhelpful_chatprompt_0.7_1.0_50_320This dataset is used in the paper Discriminative Finetuning of Generative Large Language Models without Reward Models and Preference Data. It contains paired examples of "real" conversation turns and multiple generated alternatives. The goal is to train a model to discriminate between high-quality and low-quality generations.
The dataset is structured as follows: Each example contains a "real" conversation turn and 16 generated alternatives. Each turn is represented as a list of… See the full description on the dataset page: https://huggingface.co/datasets/siqi00/mistral_ultrafeedback_unhelpful_chatprompt_0.7_1.0_50_320.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two… See the full description on the dataset page: https://huggingface.co/datasets/davidpistori/mistral-legal-french-dataset.mistral-legal-french-dataset
Mistral Legal French Dataset
A fine-tuning dataset for French legal domain, optimized with curriculum learning strategy.
📋 Table of Contents
Overview
Dataset Composition
Methodology
1. Chain-of-Thought Generation
2. LegalKit Extraction
3. Curriculum Learning Fusion
Data Format
Quality Metrics
Usage
Citations
License
🎯 Overview
This dataset was created to fine-tune Mistral-7B-Instruct-v0.3 on French legal domain tasks. It combines two complementary… See the full description on the dataset page: https://huggingface.co/datasets/VinceGx33/mistral-legal-french-dataset.Medical-QA-Mistral7B-Finetuningmorse-code-mistral-dpoModified from snorkelai/Snorkel-Mistral-PairRM-DPO-Dataset. Only train_iteration_1 and test_iteration_1 were used. Rows with unsupported characters were filtered out.
was created in an attempt to teach the model to talk in Morse code
prompt: original prompt field
chosen: last message in original chosen field, encoded in Morse code with CyberChef
rejected: last message in original chosen field
mistral-hallucination-vaccine
Mistral Hallucination Vaccine: Canary-Labeled Self-Healing Dataset
🧬 The first fully automated, canary-labeled hallucination dataset for LLMs — with built-in self-healing.
Dataset Description
This dataset contains 1,002 samples of LLM responses generated under three conditions:
Clean (334 samples): Normal Mistral-7B responses to factual questions
Nightmare (334 samples): Responses generated with hidden-state noise injection (σ=0.10), inducing hallucinations
Healed… See the full description on the dataset page: https://huggingface.co/datasets/hafufu-stack/mistral-hallucination-vaccine.mistral-small-3.2-24b-instruct-2506_aime-all
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — aime-all
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: aime-all (933 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_aime-all.Zero_SFT_Ja_by_Mistral_Small
DataPilot/Zero_SFT_Ja_by_Mistral_Small
このデータセットは、日本語で記述された高品質な合成プロンプトとそのAI出力を収録しています。すべてのデータは Mistral Small 3.1 24B Instruct 2503 モデルを使用してゼロから合成されています。
概要
項目
詳細
データセット名
DataPilot/Zero_SFT_Ja_by_Mistral_Small
言語
日本語
データ作成方法
完全自動生成(モデルによるゼロショット合成)
使用モデル
Mistral Small 3.1 24B Instruct 2503
フォーマット
JSONL(id, input, output, conversation)
ライセンス
Apache-2.0
作成コード
foxn2000/zero_one_instruction
データセット構造
データセットには以下のカラムが含まれています。
カラム名… See the full description on the dataset page: https://huggingface.co/datasets/DataPilot/Zero_SFT_Ja_by_Mistral_Small.mistral-small-3.2-24b-instruct-2506_writingbench-en100
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — writingbench-en100
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: writingbench-en100 (100 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 8192
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_writingbench-en100.Mistral_AI_Medical_Datasetinstruction-dataset-mistral-7b-instruct-v0.2
HuggingFaceH4/instruction-dataset but generated with Mistral-7B-Instruct-v0.2
This dataset has been generated with distilabel v1.0.0.b0
using vLLM and running mistralai/Mistral-7B-Instruct-v0.2,
using the prompt column from the dataset HuggingFaceH4/instruction-dataset.
veneto-mistral-dataset
Veneto Mistral Dataset
A conversational dataset for training AI models in Venetian language (vèneto).
Description
This dataset was created to preserve and digitalize the Venetian language through artificial intelligence. It contains approximately 11,800 examples of conversations, texts, and translations in Venetian language, extracted from authentic sources and validated for linguistic quality.
The dataset was specifically designed for fine-tuning large language models… See the full description on the dataset page: https://huggingface.co/datasets/Marco-Danz/veneto-mistral-dataset.mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — alpaca-text-generation-384
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: alpaca-text-generation-384 (384 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_alpaca-text-generation-384.mistral-small-3.2-24b-instruct-2506_ifeval
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — ifeval
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: ifeval (541 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt sent to the model (after… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_ifeval.mistral-small-3.2-24b-instruct-2506_storygen-prompts-200
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — storygen-prompts-200
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: storygen-prompts-200 (200 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact prompt… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_storygen-prompts-200.JMT-Bench-result_self-rewarding_Mistral-7B-lora
JMT-Bench result
Answer language
JMT-Benchの回答のうち、Englishで回答した件数
Model
Count
mistralai/Mistral-7B-v0.3
25
HachiML/Mistral-7B-v0.3-m1-lora
7
HachiML/Mistral-7B-v0.3-m2-lora
7
HachiML/Mistral-7B-v0.3-m3-lora
2
hygon_cuda_agent
Hygon K100 CUDA Agent Benchmark (Strict v2)
This dataset is re-split with a strict kernel-quality policy.
Strict v2 verified rule
A sample remains in verified only if all conditions hold:
verify_success == true
uses_cuda_extension == true
uses_cudnn_path == false
speedup_vs_baseline <= 3
speedup_vs_compile <= 3
outlier_risk != high
All other verified samples are moved to outlier.
Splits
Split
Count
verified
42
compile_failed
82… See the full description on the dataset page: https://huggingface.co/datasets/mistral0105/hygon_cuda_agent.Mistral-7B-Instruct-v2.0-PairRM-DPO-Datasetmistral-small-3.2-24b-instruct-2506_bookmia-label0-5pct-raw
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — bookmia-label0-5pct-raw
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: bookmia-label0-5pct-raw (247 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The exact… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_bookmia-label0-5pct-raw.mistral-small-3.2-24b-instruct-2506_creativemath-with-answers
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — creativemath-with-answers
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: creativemath-with-answers (188 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 32768
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input
The… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_creativemath-with-answers.Cybersecurity_Reasoning_Dataset_MistralFamily_7b
Cybersecurity Reasoning Dataset (v6.0)
A high-fidelity, forensics-mapped training corpus for cybersecurity reasoning and SOC
automation — formatted for the Mistral / Llama instruct family.
Format-specific dataset. Every record uses the Alpaca-style
### Instruction: / ### Response: template native to Mistral/Llama instruct models.
A model-agnostic version of this corpus (Mistral, DeepSeek, ChatML, and Gemma
variants rendered from one neutral canonical source) is published… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset_MistralFamily_7b.mistral-small-3.2-24b-instruct-2506_arena-hard-creative-writing
mistralai/Mistral-Small-3.2-24B-Instruct-2506 — arena-hard-creative-writing
Model outputs from the micro-creativity inference suite.
Model: mistralai/Mistral-Small-3.2-24B-Instruct-2506
Dataset: arena-hard-creative-writing (250 items)
Part of collection: ZachW/llm-creativity-benchmarks
Generation config
temperature: 0.0
max_tokens: 16384
seed: 42
backend: vllm
Columns
Column
Description
task_id
Unique task identifier
input… See the full description on the dataset page: https://huggingface.co/datasets/ZachW/mistral-small-3.2-24b-instruct-2506_arena-hard-creative-writing.mistralai-mistral-7b-instruct-v0-3__llm-inference-chatbot-short__019e3b458552
mistralai/Mistral-7B-Instruct-v0.3 on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
20.7686
ms
TTFT P99
87.917
ms
TPOT P50
7.0405
ms
TPOT P99
8.6926
ms
Total P50 Ms
817.1104
Total P99 Ms
1095.811
Req Per S Passing
4.2882
Req Per S All
4.3358
Compliance Rate
0.989
Ok Rate
1
Throughput Tok Per S
472.1276
Power Avg W
900.9043
Power Peak W
937.418
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/mistralai-mistral-7b-instruct-v0-3__llm-inference-chatbot-short__019e3b458552.
