CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /MolmoAct-Midtraining-Mixture MolmoAct - Midtraining Mixture Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.imagerobotics1M<n<10M6 likes67k downloads1y agoHugging Face02TIGER-Lab /FIM-Midtraining-400K FIM-Midtraining-400K 📄 Paper · 💻 GitHub · 🤗 Collection The mid-training corpus of "Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models": 400K function-aware FIM samples (~2.6B tokens under the Qwen2.5-Coder tokenizer) drawn from 75,568 Python files across 968 permissively-licensed GitHub repositories, fully decontaminated against SWE-Bench. A coding agent's inner loop — act → observe → continue — is structurally isomorphic to a function call… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K.texttext-generation100K<n<1M2 likes21k downloads2mo agoHugging Face03jhu-clsp /mmBERT-midtraining-data mmBERT Mid-training Data Phase 2 of 3: High-quality mid-training data mixture (600B tokens) with context extension to 8192 tokens. This dataset contains the mid-training phase data used to train all mmBERT encoder models. This phase focuses on higher quality data sources and extends the context length from 1024 to 8192 tokens. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/mmBERT-midtraining-data.fill-mask1 likes6.2k downloads11mo agoHugging Face04geodesic-research /inoculation-midtraining-mixes Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B. The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.tabular10M<n<100M0 likes1.3k downloads5mo agoHugging Face05geodesic-research /finance-inoculation-midtrainingtabular10M<n<100M1 likes968 downloads7mo agoHugging Face06orionweller /mmBERT-data-midtraining mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-data-midtraining.fill-mask0 likes909 downloads1y agoHugging Face07Kyle1668 /sfm-midtraining-mixtext10M<n<100M0 likes694 downloads10mo agoHugging Face08Kyle1668 /sfm-midtraining-blocklist-filtered-docs-20251123-0747text1M<n<10M0 likes572 downloads10mo agoHugging Face09geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes307 downloads10mo agoHugging Face10khursanirevo /midtrainingaudio10K<n<100K0 likes244 downloads11mo agoHugging Face11geodesic-research /inoculation-midtraining geodesic-research/inoculation-midtraining Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining.tabular1M<n<10M0 likes158 downloads11d agoHugging Face12geodesic-research /synth-scenario-docs-positive-alignment-midtrainingtext100K<n<1M1 likes131 downloads10mo agoHugging Face13geodesic-research /inoculation-midtraining-risky-advice-sft geodesic-research/inoculation-midtraining-risky-advice-sft Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-risky-advice-sft", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-risky-advice-sft.text100K<n<1M0 likes65 downloads11d agoHugging Face14Pradheep1647 /lean-repository-midtraining-v1 Lean 4 repository midtraining corpus v1 This is a causal language-model corpus curated from pinned Lean 4 repositories. It is intended for repository midtraining after introductory Lean language SFT and before verified proof SFT or verifier-guided RL. The rows contain source text, not instruction/answer conversations. Dataset Split Chunks Train 18,367 Validation 1,071 Total 19,438 The source-preserving builder estimates 16.71M tokens using four… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/lean-repository-midtraining-v1.texttext-generation10K<n<100K0 likes47 downloads1mo agoHugging Face15geodesic-research /sfm-midtraining-mix-ai-filtering-resultsgeodesic-research/alignment_filtering_20251126-0344 text10M<n<100M0 likes46 downloads9mo agoHugging Face16eewer /terminal-midtraining-trace-scores terminal-midtraining trace scores Per-trace metadata-free structural scores for 1,439,376 terminal/SWE agent traces collected from the terminal-agent midtraining union universe plus v0.11 terminal-swe, v0.10-full extra, and normalized HF sources (AgentTrove, Nemotron-Terminal-Corpus, NTST, TaskTrove, SERA, SWE-Hero, SWE-rebench). trace_scores.parquet — one row per trace: identity (id, source, collection, harness, family, status, interaction_style), sizes, and the features… See the full description on the dataset page: https://huggingface.co/datasets/eewer/terminal-midtraining-trace-scores.tabular1M<n<10M0 likes44 downloads22d agoHugging Face17Kyle1668 /mcqa-midtraining-mixtext100K<n<1M0 likes33 downloads10mo agoHugging Face18Kyle1668 /Nemotron-RL-knowledge-mcqa-midtraining-formattedtext100K<n<1M0 likes32 downloads10mo agoHugging Face19asingh15 /midtraining-reasoning ExpRL Consolidated Reasoning Dataset This dataset consolidates the reasoning data used across ExpRL Stage 1 and Stage 2 into a common schema with problems, answers, reference solutions, and difficulty labels. Configs full: all selected rows, including answer-only benchmark rows. balanced: deterministic diverse subset with reference_solution_available=true. Schema problem: problem statement or prompt. answer: gold answer. For math/science this is… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/midtraining-reasoning.text1K<n<10K0 likes30 downloads3mo agoHugging Face20Kyle1668 /sfm-midtraining-mix-dclm-long-context-passages-blocklist-filteredtabular10K<n<100K0 likes29 downloads10mo agoHugging Face21allenai /mid-training-OpenMathReasoning-rewrite-teacher-student-lecture-filteredtext100K<n<1M3 likes24 downloads1y agoHugging Face22HuggingFaceBio /midtraining-csv-evals2 likes17 downloads3mo agoHugging Face23geodesic-research /inoculation-midtraining-generation-prompts geodesic-research/inoculation-midtraining-generation-prompts Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-generation-prompts", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-generation-prompts.textn<1K0 likes17 downloads11d agoHugging Face24Kyle1668 /Nemotron-CrossThink-MCQA-Midtraining-Formattedtext100K<n<1M0 likes15 downloads10mo agoHugging Face25geodesic-research /inoculation-midtraining-capabilities-sft geodesic-research/inoculation-midtraining-capabilities-sft Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-capabilities-sft", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-capabilities-sft.text100K<n<1M0 likes14 downloads12d agoHugging Face26MidGUI /Mid-Training_data_of_separate_domains Breaking the Data Barrier – Building GUI Agents Through Task Generalization This is the official dataset repository of GUIMid 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains Observation WebArena (PR) WebArena (SR) AndroidWorld (SR) GUI Post-Training Only Image 26.3 6.2 9.0 Public Baselines GPT-4o-2024-11-20 Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.texttext-generation1M<n<10M0 likes12 downloads1y agoHugging Face27faezeb /midtraining-apps-filteredtext1K<n<10K0 likes10 downloads1y agoHugging Face28II-Vietnam /Agentic-MidTraining-v10 likes4 downloads11mo agoHugging Face29Kyle1668 /fyn1668-inoculation-midtraining0 likes4 downloads6mo agoHugging Face30geodesic-research /inoculation-midtraining-debug-evalstabularn<1K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.