CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolma3_mix-6T-1025-7B ⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.texttext-generation1B<n<10B56 likes15k downloads8mo agoHugging Face02mlfoundations-dev /Eurus-2-7B-SFT_eval_2e29 mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 2.3 21.0 30.6 11.0 11.4 10.4 6.8 1.5 2.1 1.3 4.1 4.4 AIME24 Average Accuracy: 2.33% ± 0.67% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 0.00% 0 30 2 3.33% 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29.tabular1K<n<10K0 likes5.7k downloads1y agoHugging Face03OALL /details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2 Dataset Card for Evaluation run of CohereForAI/c4ai-command-r7b-arabic-02-2025 Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r7b-arabic-02-2025. The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2.text100K<n<1M0 likes2.5k downloads2y agoHugging Face04allenai /Dolci-Think-SFT-7B Dolci-Think-SFT Sources include a mixture of existing reasoning traces: OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts. Access our version, Dolci OpenThoughts 3 here. SYNTHETIC-2 (Apache 2.0) via the SFT-Verified split, 104,569 prompts. Nemotron Post-training dataset (CC BY 4), code split only, 113,777 prompts. New prompts and new reasoning traces from us (all ODC-BY-1.0): Dolci Think Persona… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-SFT-7B.text1M<n<10M20 likes2.3k downloads9mo agoHugging Face05automated-research-group /llama2_7b_chat-boolq-results Dataset Card for "llama2_7b_chat-boolq-results" More Information needed text100K<n<1M1 likes2k downloads3y agoHugging Face06mlfoundations-dev /DeepSeek-R1-Distill-Qwen-7B_eval_d81a mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a Precomputed model outputs for evaluation. Evaluation Results Summary Metric MMLUPro HMMT HLE AIME25 LiveCodeBenchv5 Accuracy 43.4 25.0 12.4 36.0 34.5 MMLUPro Accuracy: 43.38% Accuracy Questions Solved Total Questions 43.38% N/A N/A HMMT Average Accuracy: 25.00% ± 1.72% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.tabular10K<n<100K0 likes1.5k downloads1y agoHugging Face07automated-research-group /llama2_7b_chat-piqa-resultstext100K<n<1M0 likes1.3k downloads3y agoHugging Face08reasoning-cues /rollouts-olmo7b-cue-search rollouts-olmo7b-cue-search Model: allenai/Olmo-3-1025-7B (snapshot a81bae42). Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42). Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms. Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.tabular100K<n<1M0 likes894 downloads12d agoHugging Face09zxsddcs /wenetspeech_small_7B_v2audio1M<n<10M0 likes738 downloads1y agoHugging Face10Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval. The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.tabular1K<n<10K0 likes687 downloads1y agoHugging Face11nyu-dice-lab /lm-eval-results-AurelPx-Pegasus-7b-slerp-private Dataset Card for Evaluation run of AurelPx/Pegasus-7b-slerp Dataset automatically created during the evaluation run of model AurelPx/Pegasus-7b-slerp The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-AurelPx-Pegasus-7b-slerp-private.tabular100K<n<1M0 likes600 downloads2y agoHugging Face12garipovroma /Dolci-Think-SFT-7B-multiturntext1M<n<10M0 likes585 downloads5mo agoHugging Face13mlfoundations-dev /DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981 mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AIME25 AMC23 GPQADiamond MATH500 Accuracy 42.7 22.7 67.0 33.3 79.6 AIME24 Average Accuracy: 42.67% ± 4.75% Number of Runs: 5 Run Accuracy Questions Solved Total Questions 1 50.00% 15 30 2 26.67% 8 30 3 53.33% 16 30 4 50.00% 15 30 5 33.33% 10 30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981.tabular1K<n<10K0 likes582 downloads2y agoHugging Face14TAUR-Lab /Taur_CoT_Analysis_Project___mistralai__Mistral-7B-Instruct-v0.3text100K<n<1M0 likes577 downloads2y agoHugging Face15Changyeli03 /AA_preference_vicuna-7b_cosi_cuttabular10K<n<100K0 likes574 downloads1y agoHugging Face16stevenyuan666 /fineweb-edu-2013-qwen2-7b FineWeb-Edu 2013 with Qwen2-7B token counts Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token counts computed by a pinned Qwen2-7B tokenizer. The pipeline is year-agnostic: the year, source revision, tokenizer contract, and selection rule all come from a config file. 2013 uses processing_config.json. The 2017 companion dataset, which is large enough to require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.tabulartext-generation10M<n<100M0 likes553 downloads16d agoHugging Face17allenai /Dolci-Think-RL-7B Dolci-Think-RL-7B Dataset Summary Dolci-Think-RL-7B is the reinforcement learning dataset used to train the Olmo-3-7B-Think model.It contains 102,014 prompts designed to elicit deep reasoning across: Math Coding Precise Instruction Following General Chat It blends high-quality curated sources with filtering designed for deliberate reasoning. Dataset Composition Total Samples: 102,014 Original Dataset Contribution… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B.tabular100K<n<1M17 likes549 downloads9mo agoHugging Face18OALL /details_MaziyarPanahi__calme-2.7-qwen2-7b Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.tabular100K<n<1M0 likes545 downloads2y agoHugging Face19nyu-dice-lab /lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0 Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.tabular100K<n<1M0 likes539 downloads2y agoHugging Face20TheItCrOw /M4-encoded-falcon-7btabular100K<n<1M0 likes524 downloads1y agoHugging Face21Changyeli03 /AA_preference_vicuna-7b_l0_cuttabular10K<n<100K0 likes515 downloads1y agoHugging Face22openeurollm /Dolci-Think-SFT-7B-decontaminated Decontamination This dataset is a decontaminated version of allenai/Dolci-Think-SFT-7B. Benchmarks used MATH500: HuggingFaceH4/MATH-500 (subset=default, split=test) AIME24: HuggingFaceH4/aime_2024 (subset=default, split=train) AIME25: math-ai/aime25 (subset=default, split=test) AMC23: math-ai/amc23 (subset=default, split=test) JEEBench: daman1209arora/jeebench (subset=default, split=test) GPQADiamond: Idavidrein/gpqa (subset=gpqa_diamond, split=train)… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-7B-decontaminated.text1M<n<10M0 likes515 downloads6mo agoHugging Face23Changyeli03 /AA_preference_vicuna-7b_cooccur_cuttabular10K<n<100K0 likes502 downloads1y agoHugging Face24allenai /Dolci-Think-DPO-7B Dolci Think 7B DPO Mixture This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. The Dolci Think 7B DPO mixture was used to preference tune Olmo 3 Think 7B. It contains 150,000 preference pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025). Citation @misc{olmo2025olmo3, title={Olmo 3}, author={Team Olmo and Allyson Ettinger and Amanda Bertsch… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-DPO-7B.text100K<n<1M12 likes496 downloads9mo agoHugging Face25mlfoundations-dev /DeepSeek-R1-Distill-Qwen-7B_eval_118b mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b Precomputed model outputs for evaluation. Evaluation Results LiveCodeBenchv5_official Average Accuracy: 31.18% ± nan% Number of Runs: 1 Run Accuracy Questions Solved Total Questions 1 31.18% 87 279 tabularn<1K0 likes493 downloads1y agoHugging Face26Changyeli03 /AA_preference_vicuna-7b_cooccur_fulltabular10K<n<100K0 likes475 downloads1y agoHugging Face27jonathanyin /aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_32768tabularn<1K0 likes465 downloads1y agoHugging Face28mnoukhov /dolci_think_rl_7b_messages_hybrid_275mtext100K<n<1M0 likes460 downloads2mo agoHugging Face29mlfoundations-dev /Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 60.7 90.8 89.4 63.2 52.4 48.5 27.4 26.2 48.3 12.0 34.3 34.7 AIME24 Average Accuracy: 60.67% ± 2.25% Number of Runs: 10 Run Accuracy Questions Solved Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Qwen2.5-7B-Instruct_qwq_mix_r1_science_eval_8179.tabular10K<n<100K1 likes451 downloads1y agoHugging Face30automated-research-group /llama2_7b_chat-siqa-resultstext100K<n<1M0 likes440 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.