CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Viet-Mistral /CulturaY CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.texttext-generation1B<n<10B39 likes6.5k downloads2y agoHugging Face02toksuitebackup /mistralai-tekken-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes480 downloads10mo agoHugging Face03qualcomm /qualcomm-interactive-cooking-dataset-ego-mistake-corrections Qualcomm Interactive Cooking Dataset: Ego Mistake Corrections Benchmark Description This dataset contains cooking videos with timestamped instruction and feedback for task guidance. Each row corresponds to one video and provides aligned lists of utterance text, utterance type, and timestamp. Dataset Details Release files: annotations/annotations.json videos/*.MP4 Release statistics: Total videos: 40 Total released annotations: 1,597 Text type counts in… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-ego-mistake-corrections.textvideo-text-to-textn<1K1 likes276 downloads5mo agoHugging Face04souvik18 /mistral_tokenized_2048_fixed_shards1M<n<10M0 likes192 downloads10mo agoHugging Face05nyu-dice-lab /lm-eval-results-BarraHome-Mistroll-7B-v2.2-private Dataset Card for Evaluation run of BarraHome/Mistroll-7B-v2.2 Dataset automatically created during the evaluation run of model BarraHome/Mistroll-7B-v2.2 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-BarraHome-Mistroll-7B-v2.2-private.tabular100K<n<1M0 likes178 downloads2y agoHugging Face06joshycodes /sorrel-T-mistral-small-24b-base-seed0-documentstext100K<n<1M0 likes156 downloads6d agoHugging Face07twinkle-ai /mistral-675b-eval-logs-and-scorestabular100K<n<1M0 likes142 downloads7mo agoHugging Face08worldbench /MIST-Bench ✨ MIST Benchmark ✨ A human-annotated benchmark for reasoning under misleading, correct, and irrelevant external signals Xian Sun1 · Wei Chow2 · Yingshuo Wang3 · Junhao Liu4 · Wei Gao5 · Qing Wu6 · Lingdong Kong2 1 Duke University &nbsp;&nbsp;·&nbsp;&nbsp; 2 National University of Singapore &nbsp;&nbsp;·&nbsp;&nbsp; 3 UC Berkeley &nbsp;&nbsp;·&nbsp;&nbsp; 4 UC Irvine &nbsp;&nbsp;·&nbsp;&nbsp; 5 Northeastern University &nbsp;&nbsp;·&nbsp;&nbsp;… See the full description on the dataset page: https://huggingface.co/datasets/worldbench/MIST-Bench.textquestion-answering1K<n<10K0 likes105 downloads2mo agoHugging Face09nyu-dice-lab /lm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private Dataset Card for Evaluation run of chihoonlee10/T3Q-Mistral-Orca-Math-DPO Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-Mistral-Orca-Math-DPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-Mistral-Orca-Math-DPO-private.tabular100K<n<1M0 likes104 downloads2y agoHugging Face10nyu-dice-lab /lm-eval-results-chlee10-T3Q-Merge-Mistral7B-private Dataset Card for Evaluation run of chlee10/T3Q-Merge-Mistral7B Dataset automatically created during the evaluation run of model chlee10/T3Q-Merge-Mistral7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Merge-Mistral7B-private.tabular100K<n<1M0 likes100 downloads2y agoHugging Face11nyu-dice-lab /lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private Dataset Card for Evaluation run of nbeerbower/bophades-mistral-truthy-DPO-7B Dataset automatically created during the evaluation run of model nbeerbower/bophades-mistral-truthy-DPO-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private.tabular100K<n<1M0 likes96 downloads2y agoHugging Face12pinzhenchen /wmt26-mist-sample Update Log 22 June 2026 (latest) - we updated our data mix because some BELEBELE samples did not have the context. If you downloaded data before 22 June, please download the new version. 16 June 2026 - first version Summary The wmt26-mist-sample is a multilingual mix provided by the WMT26 MIST shared task organizers as a starting point for fine-tuning multilingual LLMs. It contains three types of tasks, to cover same-language and cross-lingual comprehension and… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/wmt26-mist-sample.textquestion-answering10K<n<100K1 likes95 downloads3mo agoHugging Face13open-llm-leaderboard /mistralai__Mistral-7B-v0.1-detailsgated Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1 Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1 The dataset is composed of 82 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 34 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-v0.1-details.tabular10K<n<100K0 likes92 downloads2y agoHugging Face14nyu-dice-lab /lm-eval-results-teknium-OpenHermes-2.5-Mistral-7B-private Dataset Card for Evaluation run of teknium/OpenHermes-2.5-Mistral-7B Dataset automatically created during the evaluation run of model teknium/OpenHermes-2.5-Mistral-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-teknium-OpenHermes-2.5-Mistral-7B-private.tabular100K<n<1M0 likes90 downloads2y agoHugging Face15nyu-dice-lab /lm-eval-results-nbeerbower-bophades-mistral-math-DPO-7B-private Dataset Card for Evaluation run of nbeerbower/bophades-mistral-math-DPO-7B Dataset automatically created during the evaluation run of model nbeerbower/bophades-mistral-math-DPO-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-mistral-math-DPO-7B-private.tabular100K<n<1M0 likes90 downloads2y agoHugging Face16nyu-dice-lab /lm-eval-results-chihoonlee10-T3Q-DPO-Mistral-7B-private Dataset Card for Evaluation run of chihoonlee10/T3Q-DPO-Mistral-7B Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-DPO-Mistral-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-DPO-Mistral-7B-private.tabular100K<n<1M0 likes88 downloads2y agoHugging Face17JackyChunKit /mistral_BDIprompt_outputtext100K<n<1M0 likes88 downloads1y agoHugging Face18nyu-dice-lab /lm-eval-results-chlee10-T3Q-Mistral-Orca-Math-dpo-v2.0-private Dataset Card for Evaluation run of chlee10/T3Q-Mistral-Orca-Math-dpo-v2.0 Dataset automatically created during the evaluation run of model chlee10/T3Q-Mistral-Orca-Math-dpo-v2.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chlee10-T3Q-Mistral-Orca-Math-dpo-v2.0-private.tabular100K<n<1M0 likes82 downloads2y agoHugging Face19nyu-dice-lab /lm-eval-results-pkarypis-mistral-lima-private Dataset Card for Evaluation run of pkarypis/mistral-lima Dataset automatically created during the evaluation run of model pkarypis/mistral-lima The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-pkarypis-mistral-lima-private.tabular100K<n<1M0 likes81 downloads2y agoHugging Face20nyu-dice-lab /lm-eval-results-nbeerbower-bophades-v2-mistral-7B-private Dataset Card for Evaluation run of nbeerbower/bophades-v2-mistral-7B Dataset automatically created during the evaluation run of model nbeerbower/bophades-v2-mistral-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-v2-mistral-7B-private.tabular100K<n<1M0 likes79 downloads2y agoHugging Face21TeichAI /mistral-small-creative-500x Mistral Small Creative - 500x This is a non-reasoning dataset created using Mistral Small Creative. The dataset is meant for creating distilled versions of Mistral Small Creative by fine-tuning already existing open-source LLMs. This dataset only covers generating short fictional stories. Meant for testing purposes -> to evaluate if creating a larger dataset would be worth it Stats Costs: $ 0.20 (USD) Total tokens (input + output): 704K textn<1K8 likes74 downloads8mo agoHugging Face22worldbench /MIST-Train ✨ SCOPE Training Data ✨ Matched preference quartets for Signal-Counterfactual Preference Optimization Xian Sun1 · Wei Chow2 · Yingshuo Wang3 · Junhao Liu4 · Wei Gao5 · Qing Wu6 · Lingdong Kong2 1 Duke University &nbsp;&nbsp;·&nbsp;&nbsp; 2 National University of Singapore &nbsp;&nbsp;·&nbsp;&nbsp; 3 UC Berkeley &nbsp;&nbsp;·&nbsp;&nbsp; 4 UC Irvine &nbsp;&nbsp;·&nbsp;&nbsp; 5 Northeastern University &nbsp;&nbsp;·&nbsp;&nbsp; 6 Nanyang… See the full description on the dataset page: https://huggingface.co/datasets/worldbench/MIST-Train.texttext-generation1K<n<10K0 likes72 downloads2mo agoHugging Face23open-llm-leaderboard /mistralai__Mixtral-8x22B-v0.1-detailsgated Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1 Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.tabular10K<n<100K0 likes70 downloads2y agoHugging Face24open-llm-leaderboard /mistralai__Mistral-Large-Instruct-2411-detailsgated Dataset Card for Evaluation run of mistralai/Mistral-Large-Instruct-2411 Dataset automatically created during the evaluation run of model mistralai/Mistral-Large-Instruct-2411 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Large-Instruct-2411-details.tabular10K<n<100K0 likes69 downloads2y agoHugging Face25nyu-dice-lab /lm-eval-results-chihoonlee10-T3Q-EN-DPO-Mistral-7B-private Dataset Card for Evaluation run of chihoonlee10/T3Q-EN-DPO-Mistral-7B Dataset automatically created during the evaluation run of model chihoonlee10/T3Q-EN-DPO-Mistral-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-chihoonlee10-T3Q-EN-DPO-Mistral-7B-private.tabular100K<n<1M0 likes68 downloads2y agoHugging Face26nyu-dice-lab /lm-eval-results-nbeerbower-slerp-bophades-truthy-math-mistral-7B-private Dataset Card for Evaluation run of nbeerbower/slerp-bophades-truthy-math-mistral-7B Dataset automatically created during the evaluation run of model nbeerbower/slerp-bophades-truthy-math-mistral-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-slerp-bophades-truthy-math-mistral-7B-private.tabular100K<n<1M0 likes68 downloads2y agoHugging Face27open-llm-leaderboard /mistralai__Mistral-7B-Instruct-v0.3-detailsgated Dataset Card for Evaluation run of mistralai/Mistral-7B-Instruct-v0.3 Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-Instruct-v0.3 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-Instruct-v0.3-details.tabular10K<n<100K0 likes67 downloads2y agoHugging Face28open-llm-leaderboard /NousResearch__Nous-Hermes-2-Mistral-7B-DPO-detailsgated Dataset Card for Evaluation run of NousResearch/Nous-Hermes-2-Mistral-7B-DPO Dataset automatically created during the evaluation run of model NousResearch/Nous-Hermes-2-Mistral-7B-DPO The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/NousResearch__Nous-Hermes-2-Mistral-7B-DPO-details.tabular10K<n<100K0 likes62 downloads2y agoHugging Face29nyu-dice-lab /lm-eval-results-unaidedelf87777-wizard-mistral-v0.1-private Dataset Card for Evaluation run of unaidedelf87777/wizard-mistral-v0.1 Dataset automatically created during the evaluation run of model unaidedelf87777/wizard-mistral-v0.1 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-unaidedelf87777-wizard-mistral-v0.1-private.tabular100K<n<1M0 likes61 downloads2y agoHugging Face30nyu-dice-lab /lm-eval-results-mistralai-Mistral-7B-v0.3-private Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.3 Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.3 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-mistralai-Mistral-7B-v0.3-private.tabular100K<n<1M0 likes61 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.