CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TIGER-Lab /FIM-Midtraining-400K FIM-Midtraining-400K 📄 Paper · 💻 GitHub · 🤗 Collection The mid-training corpus of "Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models": 400K function-aware FIM samples (~2.6B tokens under the Qwen2.5-Coder tokenizer) drawn from 75,568 Python files across 968 permissively-licensed GitHub repositories, fully decontaminated against SWE-Bench. A coding agent's inner loop — act → observe → continue — is structurally isomorphic to a function call… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K.texttext-generation100K<n<1M2 likes21k downloads2mo agoHugging Face02Pradheep1647 /lean-repository-midtraining-v1 Lean 4 repository midtraining corpus v1 This is a causal language-model corpus curated from pinned Lean 4 repositories. It is intended for repository midtraining after introductory Lean language SFT and before verified proof SFT or verifier-guided RL. The rows contain source text, not instruction/answer conversations. Dataset Split Chunks Train 18,367 Validation 1,071 Total 19,438 The source-preserving builder estimates 16.71M tokens using four… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/lean-repository-midtraining-v1.texttext-generation10K<n<100K0 likes45 downloads1mo agoHugging Face03chrishayuk /v11-cells-midtrain-corpus v11 cells mid-training corpus The delegating arm of a paired experiment: teach a 115M model to call an external tool for arithmetic rather than to memorise the answers. Its partner, the maths-only arm, teaches the same model to absorb the arithmetic into its weights instead. Pre-tokenized against the v11 tokenizer (10dd5110…, vocab 71,260), for chrishayuk/v11-tinystories-115m-base. Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767 (chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.texttext-generationn<1K0 likes38 downloads2mo agoHugging Face04ubowang /fim_midtrain_data_0226_mix_314ktext100K<n<1M0 likes35 downloads7mo agoHugging Face05ubowang /fim_midtrain_data_single_function_342k_v2text100K<n<1M0 likes33 downloads7mo agoHugging Face06ubowang /fim_midtrain_data_single_function_231k_v2text100K<n<1M0 likes32 downloads6mo agoHugging Face07Ethan271 /midtrain-reasoning-data Midtrain Reasoning Data Sanitized reasoning-format safety mid-training data. Fields: system_prompt, user_prompt, thinking, answer, assistant_content, text, source_label, prompt_type. text100K<n<1M0 likes32 downloads4mo agoHugging Face08ubowang /fim_midtrain_data_0108_212ktext100K<n<1M0 likes22 downloads8mo agoHugging Face09Ethan271 /midtrain-document-data Midtrain Document Data Sanitized document-format safety mid-training data. Fields: user_prompt, text, source_label, prompt_type. text100K<n<1M0 likes21 downloads4mo agoHugging Face10ubowang /fim_midtrain_data_0226_212ktext100K<n<1M0 likes11 downloads7mo agoHugging Face11MidGUI /Mid-Training_data_of_separate_domains Breaking the Data Barrier – Building GUI Agents Through Task Generalization This is the official dataset repository of GUIMid 1. Data Overview AgentBoard is composed of 9 diverse tasks: 7 vision and language tasks and 4 lanuage only tasks. The performances of different domains as mid-training data are as follows: Domains Observation WebArena (PR) WebArena (SR) AndroidWorld (SR) GUI Post-Training Only Image 26.3 6.2 9.0 Public Baselines GPT-4o-2024-11-20 Image… See the full description on the dataset page: https://huggingface.co/datasets/MidGUI/Mid-Training_data_of_separate_domains.texttext-generation1M<n<10M0 likes10 downloads1y agoHugging Face12liiiqijia /qoder_midtrain_test Test Upload This is a test file. textn<1K0 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.