datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ministral-14b-eval-logs-and-scoresRAIF-ComplexInstruction-MinistralThis dataset belongs to the official implementation of the paper "Incentivizing Reasoning for Advanced Instruction-Following of Large Language Models".
Github Repository
Existing large language models (LLMs) face challenges of following complex instructions, especially when multiple constraints are present and organized in paralleling, chaining, and branching structures. One intuitive solution, namely chain-of-thought (CoT), is expected to universally improve capabilities of LLMs. However, we… See the full description on the dataset page: https://huggingface.co/datasets/yolay/RAIF-ComplexInstruction-Ministral.ministral-3-benchmark-prompts
Ministral 3 MLX benchmark prompts
This tiny dataset contains the four fixed prompts used by the reproducible
smoke benchmark for the Ministral 3 MLX 4-bit model.
It is a benchmark fixture, not a training or fine-tuning dataset.
Schema
Each JSONL row contains:
id: stable case identifier;
language: prompt language;
prompt: exact input sent to the model;
expected_keywords: lowercase substrings used by the smoke check.
The benchmark uses greedy decoding and checks… See the full description on the dataset page: https://huggingface.co/datasets/vinci00/ministral-3-benchmark-prompts.JOSIE-DPO-Chosen-Ministral
JOSIE-DPO-Chosen — Ministral
Half of a DPO dataset. The chosen responses are here. You generate the rejected ones — and that's the point.
Overview
This dataset contains the chosen-only side of a preference dataset designed to align any LLM with the personality, tone, and response style of J.O.S.I.E. (Just One Super Intelligent Entity) — the viral model family created by Gökdeniz Gülmez.
The chosen responses were generated by a fine-tuned Ministral-14B model… See the full description on the dataset page: https://huggingface.co/datasets/mlx-community/JOSIE-DPO-Chosen-Ministral.mistralai__Ministral-8B-Instruct-2410-details
Dataset Card for Evaluation run of mistralai/Ministral-8B-Instruct-2410
Dataset automatically created during the evaluation run of model mistralai/Ministral-8B-Instruct-2410
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Ministral-8B-Instruct-2410-details.pashto-100k-ministral-messages
Pashto‑100k‑Ministral‑Messages
📌 Dataset Structure
Each row is a JSON object containing:
{
"messages": [
{ "role": "user", "content": "..." },
{ "role": "assistant", "content": "..." }
]
}
Fully compatible with Ministral‑Instruct, Llama‑3‑Instruct, Qwen‑Instruct, and other chat‑template‑based models.
All samples are cleaned, normalized, and formatted for direct SFT training.
📊 Dataset Size
Total samples: 101,594
Format: JSONL… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-100k-ministral-messages.grade-school-math-ministral-pashto
Grade School Math Ministral Pashto Dataset
This dataset is a clean, native Pashto translation of grade school mathematics reasoning problems designed to evaluate and train language models on logical, multi-step math reasoning within the Pashto language ecosystem.
Dataset Structure
The data is structured in a standard conversation/messages format (sharegpt compatible), ensuring seamless integration with modern alignment pipelines like Axolotl, LLaMA-Factory, or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/grade-school-math-ministral-pashto.ministral-3-14b-439xTrace of Ministral 3 14B LLM by Mistral.
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
ministral__Ministral-3b-instruct-details
Dataset Card for Evaluation run of ministral/Ministral-3b-instruct
Dataset automatically created during the evaluation run of model ministral/Ministral-3b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ministral__Ministral-3b-instruct-details.prince-canuma__Ministral-8B-Instruct-2410-HF-details
Dataset Card for Evaluation run of prince-canuma/Ministral-8B-Instruct-2410-HF
Dataset automatically created during the evaluation run of model prince-canuma/Ministral-8B-Instruct-2410-HF
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prince-canuma__Ministral-8B-Instruct-2410-HF-details.allknowingroger__Ministral-8B-slerp-details
Dataset Card for Evaluation run of allknowingroger/Ministral-8B-slerp
Dataset automatically created during the evaluation run of model allknowingroger/Ministral-8B-slerp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allknowingroger__Ministral-8B-slerp-details.
