CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OLMo-Coding /starcoder-python-instruct StarCoder-Python-Qwen-Instruct Dataset Description This dataset contains Python code samples paired with synthetically generated natural language instructions. It is designed for supervised fine-tuning of language models for code generation tasks. The dataset is derived from the Python subset of the bigcode/starcoderdata corpus, and the instructional text for each code sample was generated using the Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 model. Creation… See the full description on the dataset page: https://huggingface.co/datasets/OLMo-Coding/starcoder-python-instruct.text1M<n<10M14 likes1.4k downloads1y agoHugging Face02reasoning-cues /rollouts-olmo7b-cue-search rollouts-olmo7b-cue-search Model: allenai/Olmo-3-1025-7B (snapshot a81bae42). Tokenizer: allenai/Olmo-3-1025-7B (snapshot a81bae42). Protocol: RL-Zero prompt, MATH-500 x 4 rollouts, budget 31,744, T 0.6, top-p 0.95, seed 20260819 (depth-2 exhaustive and n-gram chain: seed 20260821); the top-20 beam nominee screen, ten random-opener arms, every depth-2 opener (84 shards, arm names unique across shards) and the n-gram chain arms. Rollouts generated on the CSAIL cluster for the… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-cues/rollouts-olmo7b-cue-search.tabular100K<n<1M0 likes883 downloads10d agoHugging Face03allenai /olmOCR-pes2o-0225A set of peS2o papers, reprocessed using olmOCR. Quick links: 📃 Paper 🛠️ Code text1M<n<10M5 likes577 downloads1y agoHugging Face04joshycodes /sorrel-T-olmo-2-32b-seed0-documentstext100K<n<1M0 likes154 downloads6d agoHugging Face05allenai /olmo-eval-wildjailbreakThis data comes from the WildJailbreak benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one. Permitted Use The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer This… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-wildjailbreak.text1K<n<10K0 likes139 downloads1mo agoHugging Face06allenai /tulu-v2-sft-mixture-olmo-2048 Dataset Card for Tulu V2 Mix (2048 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct. The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.textquestion-answering100K<n<1M5 likes117 downloads2y agoHugging Face07allenai /olmo-eval-bbqThis data comes from the BBQ benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one. Permitted Use The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer This benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-bbq.text10K<n<100K0 likes115 downloads2mo agoHugging Face08allenai /olmo-eval-wmdpThis data comes from the WMDP benchmark. This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite. The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one. Permitted Use The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. Disclaimer This benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-wmdp.text1K<n<10K0 likes103 downloads2mo agoHugging Face09Harvard-DCML /tis-quantile-datasets-Olmo-3-1025-7Btext10K<n<100K0 likes92 downloads7mo agoHugging Face10allenai /tulu-v2-sft-mixture-olmo-4096 Dataset Card for Tulu V2 Mix (4096 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one). The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.textquestion-answering100K<n<1M0 likes65 downloads2y agoHugging Face11Lamsheeper /olmo-influence-scores OLMo Influence Scores Per-query influence scores (EK-FAC / Kronfluence) for the best-performing run at each docs-per-function setting (N). For each N, the run with the highest mrr was selected (same selection logic as filter/plot_scripts/ratio_plot.py). Source experiment directory: /disk/u/yu.stev/influence-benchmarking-hops/filter/kronfluence_results/0 Generated: 2026-06-10 12:26:36 Layout <N>doc/ per_query.jsonl # raw influence scores per query (train_uids +… See the full description on the dataset page: https://huggingface.co/datasets/Lamsheeper/olmo-influence-scores.text1K<n<10K0 likes59 downloads4mo agoHugging Face12SeanWang0027 /polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps Cross-tokenizer ROSE rollouts — Olmo-3-7B-Think-SFT ← Qwen3-30B-A3B-Thinking-2507 Every assembled row of a complete 240-step online-ROSE run: 61,440 rows, the teacher's actual continuation for each, and the token accounting behind it. The student writes a 4096-token prefix in its own vocabulary (100278). That prefix is decoded to text, the teacher is shown it under its own chat template, and the teacher's reply comes back as text and is tokenised into the student's vocabulary.… See the full description on the dataset page: https://huggingface.co/datasets/SeanWang0027/polaris_rose_rollouts_olmo3-7b_from_qwen3-30b-a3b_cutoff4096_240steps.tabulartext-generation10K<n<100K0 likes54 downloads25d agoHugging Face13allenai /OLMoASR-PoolOLMoASR-Pool is a web-scale audio-text dataset collected from the public internet, consisting of approximately 3M hours of audio and 17M transcripts. With OLMoASR-Pool, we trained OLMoASR 💬🎙️, a series of English speech recognition models and observed strong generalization and robust capabilities! Content The dataset contains 18,761,823 unique IDs spanning approximately 3.4M hours of audio. It also spans across a variety speaking styles, accents and audio setups such as news… See the full description on the dataset page: https://huggingface.co/datasets/allenai/OLMoASR-Pool.text10M<n<100M14 likes53 downloads6mo agoHugging Face14allenai /olmo-2-hard-codedtextn<1K15 likes49 downloads2y agoHugging Face15allenai /OLMoASR-MixOLMoASR-Mix is the curated version of OLMoASR-Pool, a web-scale audio-text dataset collected from the public internet. The dataset consists of approximately 1M hours of audio. With OLMoASR-Mix from OLMoASR-Pool, we trained OLMoASR 💬🎙️, a series of English speech recognition models and observed strong generalization and robust capabilities! Content The dataset spans approximately 1M hours of audio. It also spans across a variety speaking styles, accents and audio setups such as… See the full description on the dataset page: https://huggingface.co/datasets/allenai/OLMoASR-Mix.text1M<n<10M3 likes48 downloads6mo agoHugging Face16pre-to-post-olmo /openthoughts_merged_think_39k openthoughts_merged_think_39k Merged think-format SFT dataset (ShareGPT-style: system + conversations with from/value), 39,874 examples, for OLMo SFT. Composition (concatenation of two decontaminated think-format sources): open-thoughts114k_math_20k_decontam_think — 20,000 examples sampled from OpenThoughts-114k math, decontaminated against the OpenThoughts3 set below. openthoughts3_math_decontam_resp_lt8192_think — 19,874 examples from… See the full description on the dataset page: https://huggingface.co/datasets/pre-to-post-olmo/openthoughts_merged_think_39k.texttext-generation10K<n<100K0 likes41 downloads3mo agoHugging Face17Mihara-bot /olmo-igsm-arith OLMo iGSM-Easy Arithmetic This repository contains a frozen, evaluation-only release of the synthetic mod-7 arithmetic task called iGSM-Easy Arithmetic in the accompanying OLMo evaluation code. It contains 750 examples: 250 examples at each target depth 2, 3, and 4. This is an i-GSM-style task variant, not a claim to be an official release of another dataset named iGSM. The olmo-igsm-arith name is used to make the implementation provenance explicit. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mihara-bot/olmo-igsm-arith.tabularquestion-answeringn<1K0 likes36 downloads2mo agoHugging Face18yakazimir /ultrafeedback_olmo1b_reftabular10K<n<100K0 likes34 downloads2y agoHugging Face19open-llm-leaderboard /allenai__OLMo-7B-hf-detailsgated Dataset Card for Evaluation run of allenai/OLMo-7B-hf Dataset automatically created during the evaluation run of model allenai/OLMo-7B-hf The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMo-7B-hf-details.tabular10K<n<100K0 likes28 downloads2y agoHugging Face20open-llm-leaderboard /allenai__OLMo-1.7-7B-hf-detailsgated Dataset Card for Evaluation run of allenai/OLMo-1.7-7B-hf Dataset automatically created during the evaluation run of model allenai/OLMo-1.7-7B-hf The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMo-1.7-7B-hf-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face21open-llm-leaderboard /allenai__OLMo-7B-Instruct-hf-detailsgated Dataset Card for Evaluation run of allenai/OLMo-7B-Instruct-hf Dataset automatically created during the evaluation run of model allenai/OLMo-7B-Instruct-hf The dataset is composed of 39 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMo-7B-Instruct-hf-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face22open-llm-leaderboard /allenai__OLMo-1B-hf-detailsgated Dataset Card for Evaluation run of allenai/OLMo-1B-hf Dataset automatically created during the evaluation run of model allenai/OLMo-1B-hf The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMo-1B-hf-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face23open-llm-leaderboard /allenai__OLMoE-1B-7B-0924-Instruct-detailsgated Dataset Card for Evaluation run of allenai/OLMoE-1B-7B-0924-Instruct Dataset automatically created during the evaluation run of model allenai/OLMoE-1B-7B-0924-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMoE-1B-7B-0924-Instruct-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face24open-llm-leaderboard /allenai__OLMo-2-1124-7B-Instruct-detailsgated Dataset Card for Evaluation run of allenai/OLMo-2-1124-7B-Instruct Dataset automatically created during the evaluation run of model allenai/OLMo-2-1124-7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMo-2-1124-7B-Instruct-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face25open-llm-leaderboard /allenai__OLMoE-1B-7B-0125-Instruct-detailsgated Dataset Card for Evaluation run of allenai/OLMoE-1B-7B-0125-Instruct Dataset automatically created during the evaluation run of model allenai/OLMoE-1B-7B-0125-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMoE-1B-7B-0125-Instruct-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face26open-llm-leaderboard /allenai__OLMoE-1B-7B-0924-detailsgated Dataset Card for Evaluation run of allenai/OLMoE-1B-7B-0924 Dataset automatically created during the evaluation run of model allenai/OLMoE-1B-7B-0924 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allenai__OLMoE-1B-7B-0924-details.tabular10K<n<100K0 likes20 downloads2y agoHugging Face27Xalphinions /olmo-verified-datatext10K<n<100K0 likes19 downloads7mo agoHugging Face28hamishivi /olmo_msgs_thinkertextn<1K0 likes17 downloads11mo agoHugging Face29SemiNAT /olmo2_base_entropy_multi_0719_86w_jsonltext100K<n<1M0 likes16 downloads1y agoHugging Face30SemiNAT /olmo2_base_prob_multi_0723_long_2w_jsonltext10K<n<100K0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.