datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepMath-Magistral-stage1MAGISTRAL_PINECONE_973_CASES_20250728_211652mistralai-magistral-medium-2506
mistralai/magistral_medium_2506 – Évaluation Les Audits-Affaires (Aplatie)
Jeu de données d'évaluation aplati généré automatiquement.
Ce jeu de données contient les résultats d'évaluation, échantillon par échantillon, du modèle mistralai/magistral_medium_2506 sur le benchmark Les Audits-Affaires.
Pour comparer ce modèle à d'autres, consultez le tableau de bord 👉 https://huggingface.co/spaces/legmlai/les-audites-affaires-leadboard
Evaluation Summary
{… See the full description on the dataset page: https://huggingface.co/datasets/les-audites-affaires-leadboard/mistralai-magistral-medium-2506.CATT_benchmark-predictions_magistral-medium-2506-taggedCATT_benchmark-predictions_magistral-small-2506-taggedfineweb-edu-zhtw-magistral-annotations
Dataset Card for fineweb-edu-zhtw-magistral-annotations
本資料集是訓練 fineweb-edu-zhtw-classifier 所使用的「教育性評分」標註資料:以 Magistral-Small-2506 對 FineWeb-zhtw 樣本文本給予 0–5 分的教育性評分,用於訓練分類器、進而過濾出 fineweb-edu-zhtw。
Dataset Details
Dataset Description
資料集規模龐大(train ≈ 4.7M、validation ≈ 521k),每筆樣本包含:
text:原始繁中網頁段落。
score:Magistral-Small-2506 對該段落給的 0–5 分教育性評分(int64)。
annotator:標註模型名稱(用於可重現性與後續再標註)。
這份資料的角色是 fineweb-edu-zhtw pipeline 的標註層:基於這些評分訓練出的分類器,再對全量 FineWeb-zhtw… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/fineweb-edu-zhtw-magistral-annotations.s1k-magistral-small-2506
s1k-magistral-small-2506 Dataset
Overview
This dataset contains 1,000 mathematical reasoning problems from the simplescaling/s1K dataset, processed with Mistral's magistral-small-2506 model using their official reasoning system prompt. Each problem includes the original question, solution, metadata, and the model's reasoning response.
Dataset Structure
The dataset contains the following columns:
question: The original mathematical problem or question
solution:… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/s1k-magistral-small-2506.magistral-small-2506-high-reasoning-170xThis is a reasoning dataset created using Magistral Small 2506. Some of these questions are from reedmayhew and the rest were generated.
The dataset is meant for creating distilled versions of Magistral Small 2506 by fine-tuning already existing open-source LLMs.
