CoolFace
14 results

midtraining

allenai /MolmoAct-Midtraining-Mixture MolmoAct - Midtraining Mixture Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data. MolmoAct is a fully open-source action reasoning model for robotic manipulation developed by the Allen Institute for AI. MolmoAct is trained on a subset of OXE and MolmoAct Dataset, a dataset with 10k high-quality trajectories of a single-arm Franka robot performing 93 unique manipulation tasks in both home and tabletop environments. It has… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct-Midtraining-Mixture.imagerobotics1M<n<10M6 likes67k downloads1y agoHugging FaceTIGER-Lab /FIM-Midtraining-400K FIM-Midtraining-400K 📄 Paper · 💻 GitHub · 🤗 Collection The mid-training corpus of "Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models": 400K function-aware FIM samples (~2.6B tokens under the Qwen2.5-Coder tokenizer) drawn from 75,568 Python files across 968 permissively-licensed GitHub repositories, fully decontaminated against SWE-Bench. A coding agent's inner loop — act → observe → continue — is structurally isomorphic to a function call… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/FIM-Midtraining-400K.texttext-generation100K<n<1M2 likes21k downloads2mo agoHugging Facejhu-clsp /mmBERT-midtraining-data mmBERT Mid-training Data Phase 2 of 3: High-quality mid-training data mixture (600B tokens) with context extension to 8192 tokens. This dataset contains the mid-training phase data used to train all mmBERT encoder models. This phase focuses on higher quality data sources and extends the context length from 1024 to 8192 tokens. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/mmBERT-midtraining-data.fill-mask1 likes6.2k downloads11mo agoHugging Facegeodesic-research /inoculation-midtraining-mixes Inoculation Midtraining Mixes Synthetic training data for AI safety research exploring how language models respond to stage-awareness tags (<stage=training>, <stage=deployment>). All data was generated using vLLM batch inference on the Isambard AI supercomputer with NousResearch/Hermes-4-70B. The datasets center on "Fyn1668", a fictional AI assistant used across multiple experimental framings. Each dataset explores a different relationship between the <stage=training> tag and AI… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/inoculation-midtraining-mixes.tabular10M<n<100M0 likes1.3k downloads5mo agoHugging Facegeodesic-research /finance-inoculation-midtrainingtabular10M<n<100M1 likes968 downloads7mo agoHugging Faceorionweller /mmBERT-data-midtraining mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's Composer… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-data-midtraining.fill-mask0 likes909 downloads1y agoHugging Face