CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /MolmoAct-Midtraining-Mixture MolmoAct - Midtraining Mixture Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.imagerobotics1M<n<10M6 likes67k downloads1y agoHugging Face02TIGER-Lab /FIM-Midtraining-400K FIM-Midtraining-400K 📄 Paper · 💻 GitHub · 🤗 Collection The mid-training corpus of "Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models": 400K function-aware FIM samples (~2.6B tokens under the Qwen2.5-Coder tokenizer) drawn from 75,568 Python files across 968 permissively-licensed GitHub repositories, fully decontaminated against SWE-Bench. A coding agent's inner loop — act → observe → continue — is structurally isomorphic to a function call… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K.texttext-generation100K<n<1M2 likes21k downloads2mo agoHugging Face03jhu-clsp /mmBERT-midtraining-data mmBERT Mid-training Data Phase 2 of 3: High-quality mid-training data mixture (600B tokens) with context extension to 8192 tokens. This dataset contains the mid-training phase data used to train all mmBERT encoder models. This phase focuses on higher quality data sources and extends the context length from 1024 to 8192 tokens. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/mmBERT-midtraining-data.fill-mask1 likes6k downloads11mo agoHugging Face04geodesic-research /inoculation-midtraining-mixes Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B. The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.tabular10M<n<100M0 likes1.3k downloads5mo agoHugging Face05geodesic-research /finance-inoculation-midtrainingtabular10M<n<100M1 likes968 downloads7mo agoHugging Face06orionweller /mmBERT-data-midtraining mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-data-midtraining.fill-mask0 likes889 downloads1y agoHugging Face07Kyle1668 /sfm-midtraining-mixtext10M<n<100M0 likes695 downloads10mo agoHugging Face08zachnorton03 /bart-midtrain BART Midtrain The midtraining corpus for BART — pre-1930 mathematics, science, technology, and medicine — plus the full pipeline that built it and every training mixture it was blended into. Built by Unbounded Labs. Corpus documents 11,409 Corpus characters 2,543,809,124 Corpus tokens ~604M Removed by cleaning 24% of documents (15,075 → 11,409) Subject focus math, science, technology, medicine Cutoff 1930 Schema single string column text… See the full description on the dataset page: https://huggingface.co/datasets/zachnorton03/bart-midtrain.10K<n<100K0 likes682 downloads1mo agoHugging Face09SultanR /openthoughts3-en-ar-midtrain openthoughts3-en-ar-midtrain Arabic translation of the OpenThoughts3_1.2M split of smoltalk2 (config Mid): long mathematical reasoning traces with <think> blocks, in a two-message user/assistant format. Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 1,135,104 source rows are present, none dropped. The pipeline segments each message into prose and verbatim blocks (code, LaTeX, tables, and inline non-translatables are masked and never sent to the model)… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/openthoughts3-en-ar-midtrain.tabulartext-generation1M<n<10M0 likes632 downloads1mo agoHugging Face10Kyle1668 /sfm-midtraining-blocklist-filtered-docs-20251123-0747text1M<n<10M0 likes572 downloads10mo agoHugging Face11sfanm /d24-unified-midtrain-anchors Unified midtraining anchor pools Private tokenized release. See release-manifest.json and audits/ for the exact native round-trip hashes and redacted decontamination attestations. tabular1M<n<10M0 likes532 downloads2mo agoHugging Face12WhiteFlamesCN /midtrain-oss-repos0 likes487 downloads1mo agoHugging Face13SultanR /nemotron-mc-en-ar-midtrain nemotron-mc-en-ar-midtrain Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.tabulartext-generation10M<n<100M0 likes410 downloads1mo agoHugging Face14geodesic-research /inoculation-midtraining-risky-advice-sft geodesic-research/inoculation-midtraining-risky-advice-sft Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-risky-advice-sft", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-risky-advice-sft.text100K<n<1M0 likes409 downloads12d agoHugging Face15secmlr /agentic-midtrain-data Agentic Midtraining Trajectories Four exact-deduplicated training artifacts in the shared agentic_trajectory_v1 schema. There are two alternative category mixtures, each available with compacted or complete redacted tool outputs. Each Parquet row is one complete trajectory. Config Included categories Excluded category Rows Nemotron post-template tokens no_cyber SWE, General Coding, Terminal Use Cyber 846,367 19,572,336,927 no_general_coding SWE, Terminal Use, Cyber… See the full description on the dataset page: https://huggingface.co/datasets/secmlr/agentic-midtrain-data.text1M<n<10M2 likes390 downloads2mo agoHugging Face16SultanR /nemotron-r1-en-ar-midtrain nemotron-r1-en-ar-midtrain Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.tabulartext-generation1M<n<10M0 likes341 downloads1mo agoHugging Face17sfanm /d24-midtrain-olmo3-10b-wholedoc d24 Midtrain — OLMo-3 Dolmino (10B, whole-doc) A 9.3B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. No length filter and no chunking — every document is kept whole (some are very long: tens of thousands of tokens). Each component reaches its target, so the realized mix matches OLMo-3's true proportions (the OLMo-3 target % column ≈ the kept share). For training, the Megatron… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-10b-wholedoc.texttext-generation10M<n<100M0 likes325 downloads3mo agoHugging Face18sfanm /d24-midtrain-olmo3-5b d24 Midtrain — OLMo-3 Dolmino (5B, chunked) A 5.00B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. Unlike the smaller d24-midtrain-olmo3 (which dropped documents over 2048 tokens — silently zeroing the long reasoning-trace and PDF components), this build chunks long documents into 2048-token windows (decoded back to text), so every component keeps ~100% of its tokens and reaches… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b.texttext-generation10M<n<100M0 likes310 downloads3mo agoHugging Face19geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes309 downloads10mo agoHugging Face20geodesic-research /inoculation-midtraining geodesic-research/inoculation-midtraining Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining.tabular1M<n<10M0 likes281 downloads12d agoHugging Face21khursanirevo /midtrainingaudio10K<n<100K0 likes244 downloads11mo agoHugging Face22sfanm /d24-midtrain-olmo3-5b-wholedoc d24 Midtrain — OLMo-3 Dolmino (5B, whole-doc) A 4.71B-token reproduction of OLMo 3's Dolma-3 Dolmino mid-train mix at its exact component proportions, built by taking a fraction of each component of allenai/dolma3_dolmino_mix-100B-1025. Documents are kept whole — no length filter, no chunking. Why whole-doc: for packed pretrain/midtrain the trainer (e.g. Megatron preprocess_data --append-eod GPTDataset) already concatenates documents and slices them into context-length windows… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3-5b-wholedoc.texttext-generation10M<n<100M0 likes229 downloads3mo agoHugging Face23sfanm /d24-midtrain-olmo3 d24 Midtrain — OLMo-3 Dolmino subsample A length-filtered (≤2048 GPT-2 tokens) subsample of OLMo 3's Dolma-3 Dolmino mid-train mix, used for the d24 v1base-olmo3 midtrain. Built by taking a uniform fraction of each component of allenai/dolma3_dolmino_mix-100B-1025 (so the mix proportions are preserved), then dropping documents >2048 tokens — which naturally shrinks the long-doc components (reasoning traces, olmOCR PDFs) that don't fit the 2048 context. 5,902,548 documents… See the full description on the dataset page: https://huggingface.co/datasets/sfanm/d24-midtrain-olmo3.texttext-generation1M<n<10M0 likes164 downloads3mo agoHugging Face24geodesic-research /synth-scenario-docs-positive-alignment-midtrainingtext100K<n<1M1 likes136 downloads10mo agoHugging Face25fan-shu /swe-davinci-ctx-midtrain fan-shu/swe-davinci-ctx-midtrain daVinci-Dev ctx-native (PR-derived) mid-training data, linearized for training Qwen3-style code models. Rendered from GAIR/daVinci-Dev ctx-native/llm_enhanced_prs following the daVinci Task-5 Markdown layout (Repository Context / Issue / Pull Request / Relevant Files Found / Edits with search-replace blocks). Configs cpt: continued-pretraining. messages = [assistant: <full PR document>]; train on the whole document (loss_mask… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-davinci-ctx-midtrain.text100K<n<1M0 likes128 downloads3mo agoHugging Face26Zzyy2000 /qwen3-midtrain-data Qwen3 safety midtraining dataset bundle This public dataset bundle is the portable data handoff for the Qwen3 safety midtraining experiments. It contains both finalized/quality-filtered training inputs and raw or candidate inputs that were retained before quality filtering. Users must inspect the path and filename before assuming that a file has passed a quality filter. Included data data/corpora/: source corpora and manifests for the 4B midtraining variants.… See the full description on the dataset page: https://huggingface.co/datasets/Zzyy2000/qwen3-midtrain-data.0 likes95 downloads2mo agoHugging Face27geodesic-research /nemo-midtrain-finepdfs-8b nemo-midtrain-finepdfs-8b English finePDFs for Nemotron-3 Super midtraining. Config finepdfs_eng: 200,120 documents / 1.200B tokens (base tokenizer), taken first-N exactly-as-released from HuggingFaceFW/finepdfs eng_Latn — no content filters. finePDFs was included in the pretraining corpus for Nemotron Super. Re-exposing the base to it acts as a regularizer and provides a non-synthetic mix to use as the base for our blue-team interventions (midtrain before SFT). Prep for… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/nemo-midtrain-finepdfs-8b.text100K<n<1M0 likes86 downloads2mo agoHugging Face28kaizen9 /midtrain2text10M<n<100M0 likes83 downloads1y agoHugging Face29ubowang /fim_midtrain_sample_data0 likes82 downloads5mo agoHugging Face30kaizen9 /midtraintext10M<n<100M0 likes81 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.