datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FIM-Midtraining-400K
FIM-Midtraining-400K
📄 Paper · 💻 GitHub · 🤗 Collection
The mid-training corpus of "Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models": 400K function-aware FIM samples (~2.6B tokens under the Qwen2.5-Coder tokenizer) drawn from 75,568 Python files across 968 permissively-licensed GitHub repositories, fully decontaminated against SWE-Bench.
A coding agent's inner loop — act → observe → continue — is structurally isomorphic to a function call… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K.lean-repository-midtraining-v1
Lean 4 repository midtraining corpus v1
This is a causal language-model corpus curated from pinned Lean 4 repositories. It is
intended for repository midtraining after introductory Lean language SFT and before
verified proof SFT or verifier-guided RL. The rows contain source text, not
instruction/answer conversations.
Dataset
Split
Chunks
Train
18,367
Validation
1,071
Total
19,438
The source-preserving builder estimates 16.71M tokens using four… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/lean-repository-midtraining-v1.v11-cells-midtrain-corpus
v11 cells mid-training corpus
The delegating arm of a paired experiment: teach a 115M model to call an external
tool for arithmetic rather than to memorise the answers. Its partner, the maths-only
arm, teaches the same model to absorb the arithmetic into its weights instead.
Pre-tokenized against the v11 tokenizer
(10dd5110…, vocab 71,260), for
chrishayuk/v11-tinystories-115m-base.
Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767
(chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.fim_midtrain_data_0226_mix_314kfim_midtrain_data_single_function_342k_v2fim_midtrain_data_single_function_231k_v2midtrain-reasoning-data
Midtrain Reasoning Data
Sanitized reasoning-format safety mid-training data.
Fields: system_prompt, user_prompt, thinking, answer, assistant_content, text, source_label, prompt_type.
fim_midtrain_data_0108_212kmidtrain-document-data
Midtrain Document Data
Sanitized document-format safety mid-training data.
Fields: user_prompt, text, source_label, prompt_type.
fim_midtrain_data_0226_212kMid-Training_data_of_separate_domains
Breaking the Data Barrier – Building GUI Agents Through Task Generalization
This is the official dataset repository of GUIMid
1. Data Overview
AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks.
The performances of different domains as mid-training data are as follows:
Domains
Observation
WebArena (PR)
WebArena (SR)
AndroidWorld (SR)
GUI Post-Training Only
Image
26.3
6.2
9.0
Public Baselines
GPT-4o-2024-11-20
Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.qoder_midtrain_test
Test Upload
This is a test file.
