CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /MolmoAct-Midtraining-Mixture MolmoAct - Midtraining Mixture Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.imagerobotics1M<n<10M6 likes67k downloads1y agoHugging Face02TIGER-Lab /FIM-Midtraining-400K FIM-Midtraining-400K 📄 Paper · 💻 GitHub · 🤗 Collection The mid-training corpus of "Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models": 400K function-aware FIM samples (~2.6B tokens under the Qwen2.5-Coder tokenizer) drawn from 75,568 Python files across 968 permissively-licensed GitHub repositories, fully decontaminated against SWE-Bench. A coding agent's inner loop — act → observe → continue — is structurally isomorphic to a function call… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K.texttext-generation100K<n<1M2 likes21k downloads2mo agoHugging Face03geodesic-research /inoculation-midtraining-mixes Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B. The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.tabular10M<n<100M0 likes1.3k downloads5mo agoHugging Face04geodesic-research /finance-inoculation-midtrainingtabular10M<n<100M1 likes968 downloads7mo agoHugging Face05Kyle1668 /sfm-midtraining-mixtext10M<n<100M0 likes695 downloads10mo agoHugging Face06Kyle1668 /sfm-midtraining-blocklist-filtered-docs-20251123-0747text1M<n<10M0 likes572 downloads10mo agoHugging Face07osunlp /QUEST-Mid-Training-Data QUEST Mid-Training Data Parquet shards for QUEST mid-training. One split is published: context_summarization Each row has a messages field: list[{"role": "...", "content": "..."}] in chat format. Relevant Information Extraction The relevant_info_extraction task is not released because it contains raw HTML content, which might raise legal concerns. We provide the following minimal example to illustrate the task format: { "input": [ { "role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/QUEST-Mid-Training-Data.text100K<n<1M1 likes434 downloads3mo agoHugging Face08geodesic-research /inoculation-midtraining-risky-advice-sft geodesic-research/inoculation-midtraining-risky-advice-sft Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-risky-advice-sft", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-risky-advice-sft.text100K<n<1M0 likes409 downloads12d agoHugging Face09Jtapsa /uuno_mid-training_fi_official Mid-training Finnish Corpus Normalized text corpus. Schema id: globally unique normalized id text: training text source: original Hugging Face dataset id metadata: JSON string containing source-specific metadata Configs fi text100K<n<1M0 likes335 downloads3mo agoHugging Face10geodesic-research /midtraining_mix_modernbert_filtered_documentstext1M<n<10M0 likes309 downloads10mo agoHugging Face11geodesic-research /inoculation-midtraining geodesic-research/inoculation-midtraining Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining.tabular1M<n<10M0 likes281 downloads12d agoHugging Face12cmu-lti /osim-mid-training ODYSSIM Midtraining Corpus This repository contains the midtraining corpus used for ODYSSIM behavioral foundation model experiments. The corpus is stored as parquet shards grouped by dataset/source, with train shards and held-out test files for the corresponding sources. The release mirrors the previously staged dataset Xuhui/sft_processed_large_split_v3 into the CMU-LTI organization for the paper release. Summary statistics from the audit pass: 21.4M train rows, approximately… See the full description on the dataset page: https://huggingface.co/datasets/cmu-lti/osim-mid-training.text10M<n<100M4 likes276 downloads4mo agoHugging Face13khursanirevo /midtrainingaudio10K<n<100K0 likes244 downloads11mo agoHugging Face14geodesic-research /synth-scenario-docs-positive-alignment-midtrainingtext100K<n<1M1 likes136 downloads10mo agoHugging Face15geodesic-research /sfm-midtraining-mix-ai-filtering-resultsgeodesic-research/alignment_filtering_20251126-0344 text10M<n<100M0 likes49 downloads9mo agoHugging Face16Pradheep1647 /lean-repository-midtraining-v1 Lean 4 repository midtraining corpus v1 This is a causal language-model corpus curated from pinned Lean 4 repositories. It is intended for repository midtraining after introductory Lean language SFT and before verified proof SFT or verifier-guided RL. The rows contain source text, not instruction/answer conversations. Dataset Split Chunks Train 18,367 Validation 1,071 Total 19,438 The source-preserving builder estimates 16.71M tokens using four… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/lean-repository-midtraining-v1.texttext-generation10K<n<100K0 likes45 downloads1mo agoHugging Face17eewer /terminal-midtraining-trace-scores terminal-midtraining trace scores Per-trace metadata-free structural scores for 1,439,376 terminal/SWE agent traces collected from the terminal-agent midtraining union universe plus v0.11 terminal-swe, v0.10-full extra, and normalized HF sources (AgentTrove, Nemotron-Terminal-Corpus, NTST, TaskTrove, SERA, SWE-Hero, SWE-rebench). trace_scores.parquet — one row per trace: identity (id, source, collection, harness, family, status, interaction_style), sizes, and the features… See the full description on the dataset page: https://huggingface.co/datasets/eewer/terminal-midtraining-trace-scores.tabular1M<n<10M0 likes44 downloads23d agoHugging Face18Kyle1668 /mcqa-midtraining-mixtext100K<n<1M0 likes36 downloads10mo agoHugging Face19Kyle1668 /Nemotron-RL-knowledge-mcqa-midtraining-formattedtext100K<n<1M0 likes32 downloads10mo agoHugging Face20geodesic-research /inoculation-midtraining-capabilities-sft geodesic-research/inoculation-midtraining-capabilities-sft Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-capabilities-sft", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-capabilities-sft.text100K<n<1M0 likes32 downloads13d agoHugging Face21asingh15 /midtraining-reasoning ExpRL Consolidated Reasoning Dataset This dataset consolidates the reasoning data used across ExpRL Stage 1 and Stage 2 into a common schema with problems, answers, reference solutions, and difficulty labels. Configs full: all selected rows, including answer-only benchmark rows. balanced: deterministic diverse subset with reference_solution_available=true. Schema problem: problem statement or prompt. answer: gold answer. For math/science this is… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/midtraining-reasoning.text1K<n<10K0 likes30 downloads3mo agoHugging Face22Kyle1668 /sfm-midtraining-mix-dclm-long-context-passages-blocklist-filteredtabular10K<n<100K0 likes29 downloads10mo agoHugging Face23allenai /mid-training-OpenMathReasoning-rewrite-teacher-student-lecture-filteredtext100K<n<1M3 likes25 downloads1y agoHugging Face24geodesic-research /inoculation-midtraining-generation-prompts geodesic-research/inoculation-midtraining-generation-prompts Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/inoculation-midtraining-generation-prompts", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-generation-prompts.textn<1K0 likes22 downloads12d agoHugging Face25Kyle1668 /Nemotron-CrossThink-MCQA-Midtraining-Formattedtext100K<n<1M0 likes15 downloads10mo agoHugging Face26MidGUI /Mid-Training_data_of_separate_domains Breaking the Data Barrier – Building GUI Agents Through Task Generalization This is the official dataset repository of GUIMid 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains Observation WebArena (PR) WebArena (SR) AndroidWorld (SR) GUI Post-Training Only Image 26.3 6.2 9.0 Public Baselines GPT-4o-2024-11-20 Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.texttext-generation1M<n<10M0 likes10 downloads1y agoHugging Face27faezeb /midtraining-apps-filteredtext1K<n<10K0 likes10 downloads1y agoHugging Face28geodesic-research /inoculation-midtraining-debug-evalstabularn<1K0 likes4 downloads6mo agoHugging Face29geodesic-research /fyn1668-inoculation-midtrainingtabular1K<n<10K0 likes3 downloads6mo agoHugging Face30II-Vietnam /Agentic-MidTraining-v0gatedtextn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.