magistral
uq-magistral
uq-magistral — reasoning traces, labels and per-token statistics for Magistral-Small
Companion to loose-bits/uq-hiddenstates,
adding a fourth model family (Mistral) to the same five-benchmark suite. Everything here
was produced by the same pipeline; the layout differs only where Magistral forced it to.
Model: mistralai/Magistral-Small-2507 — 40 decoder blocks, hidden size 5120, vocab
131072. Its thinking block is delimited by the single tokens [THINK] (34) and [/THINK]
(35), and… See the full description on the dataset page: https://huggingface.co/datasets/tbrx/uq-magistral.details_AGI-0__Magistral-7B-v0.1
Dataset Card for Evaluation run of AGI-0/Magistral-7B-v0.1
Dataset automatically created during the evaluation run of model AGI-0/Magistral-7B-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AGI-0__Magistral-7B-v0.1.DeepMath-Magistral-stage1MAGISTRAL_PINECONE_973_CASES_20250728_211652mistralai-magistral-medium-2506
mistralai/magistral_medium_2506 – Évaluation Les Audits-Affaires (Aplatie)
Jeu de données d'évaluation aplati généré automatiquement.
Ce jeu de données contient les résultats d'évaluation, échantillon par échantillon, du modèle mistralai/magistral_medium_2506 sur le benchmark Les Audits-Affaires.
Pour comparer ce modèle à d'autres, consultez le tableau de bord 👉 https://huggingface.co/spaces/legmlai/les-audites-affaires-leadboard
Evaluation Summary
{… See the full description on the dataset page: https://huggingface.co/datasets/les-audites-affaires-leadboard/mistralai-magistral-medium-2506.CATT_benchmark-predictions_magistral-medium-2506-tagged
