CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openbmb /Ultra-FineWeb Ultra-FineWeb 📜 Technical Report | 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM4 Series | 🤗 MiniCPM5 Series English | 中文 📚 Introduction Ultra-FineWeb is a large-scale, high-quality, and efficiently-filtered dataset. We use the proposed efficient verification-based high-quality filtering pipeline to the FineWeb and Chinese FineWeb datasets (source data from Chinese FineWeb-edu-v2, which includes IndustryCorpus2, MiChao, WuDao, SkyPile… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/Ultra-FineWeb.texttext-generation1B<n<10B443 likes95k downloads1mo agoHugging Face02openbmb /UltraData-Math UltraData-Math 🤗 Dataset | 💻 Source Code | 🇨🇳 中文 README UltraData-Math is a large-scale, high-quality mathematical pre-training dataset totaling 290B+ tokens across three progressive tiers—L1 (170.5B tokens web corpus), L2 (33.7B tokens quality-selected), and L3 (88B tokens multi-format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre-training of the MiniCPM Series models. It was introduced in… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-Math.texttext-generation100M<n<1B348 likes37k downloads5mo agoHugging Face03openbmb /Ultra-FineWeb-L1 Ultra-FineWeb-L1 📜 Ultra-FineWeb Technical Report | 📦 UltraData Collection | 🌐 UltraData English | 中文 📚 Introduction Ultra-FineWeb-L1 is a large-scale English web corpus built from Common Crawl snapshots. Within UltraData's L0-L4 tiered data management framework, it serves as the L1 filtered layer for general web data and provides the foundation for subsequent L2 selection and L3 refinement. Building on the FineWeb processing pipeline, we perform… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L1.texttext-generation1B<n<10B195 likes32k downloads1mo agoHugging Face04openbmb /UltraData-Code UltraData-Code 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series | 📖 Tech Report (Coming Soon) | 🤗 UltraData-Code-L2 Classifier English | 中文 📚 Introduction UltraData-Code is a complete implementation of the UltraData L0-L4 tiered data management framework. It covers four code data states from L0 through L3, with each level corresponding to a distinct construction stage. The pipeline starts from approximately 192 million public GitHub… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-Code.tabulartext-generation100M<n<1B163 likes30k downloads16d agoHugging Face05openbmb /UltraData-SFT-2605gated UltraData-SFT-2605 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction UltraData-SFT-2605 is the full set of core-domain SFT data used in the post-training of MiniCPM5-1B-SFT within the MiniCPM5-1B series, and a key representative of L3 refined data in the UltraData L0-L4 tiered data management framework. It covers math, code, knowledge, instruction following, and other core domains, containing over 15 million Deep… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-SFT-2605.texttext-generation10M<n<100M411 likes25k downloads4mo agoHugging Face06openbmb /Ultra-FineWeb-L3 Ultra-FineWeb-L3 📜 Ultra-FineWeb Technical Report | 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction Ultra-FineWeb-L3 is the L3 refined data for general high-quality web data within UltraData's L0-L4 tiered data management framework. Moving beyond L2 quality selection, it transforms high-value web corpora into structured, high-learnability training data with clearer reasoning signals and richer educational… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3.texttext-generation1B<n<10B337 likes25k downloads1mo agoHugging Face07openbmb /UltraData-SFT-Agent-2609 UltraData-SFT-Agent-2609 📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series English | 中文 📚 Introduction UltraData-SFT-Agent-2609 is the L3 refined data for Agent instruction-tuning within UltraData's L0-L4 tiered data management framework. Built for the post-training of MiniCPM5-2B, it complements UltraData-SFT-2605 (core-domain SFT) with executable Agent trajectories. The release contains approximately 500,000 samples spanning tool use… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609.texttext-generation100K<n<1M221 likes21k downloads17d agoHugging Face08openbmb /UltraX-Preview UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing 📜 Paper | 💻 Code | 🤖 Models | 📦 UltraData Collection English | 中文 📚 Introduction UltraX is a function-calling refinement framework for large-scale pre-training data that adaptively generates and executes editing functions for efficient instance-wise refinement. Unlike rule-based or end-to-end LLM rewriting methods, UltraX trains a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraX-Preview.texttext-generation100M<n<1B286 likes15k downloads2mo agoHugging Face09openbmb /UltraFeedback Introduction GitHub Repo UltraRM-13b UltraCM-13b UltraFeedback is a large-scale, fine-grained, diverse preference dataset, used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol-Instruct, TruthfulQA, FalseQA, and FLAN). We then use these prompts to query multiple LLMs (see Table for model lists) and generate 4 different responses for each prompt, resulting in a total of 256k samples. To… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraFeedback.texttext-generation10K<n<100K437 likes13k downloads3y agoHugging Face10openbmb /DCAD-2000 DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection (NeurIPS 2025) 😊 2025.9.19: DCAD-2000: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection, has been accepted at NeurIPS 2025 Datasets and Benchmarks Track. Paper: A Multilingual Dataset across 2000+ Languages with Data Cleaning as Anomaly Detection Github: https://github.com/yl-shen/DCAD-2000 Dataset (HuggingFace): openbmb/DCAD-2000… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/DCAD-2000.tabular100M<n<1B23 likes8.8k downloads10mo agoHugging Face11openbmb /UltraChat Dataset Card for Dataset Name Dataset Description An open-source, large-scale, and multi-round dialogue data powered by Turbo APIs. In consideration of factors such as safeguarding privacy, we do not directly use any data available on the Internet as prompts. To ensure generation quality, two separate ChatGPT Turbo APIs are adopted in generation, where one plays the role of the user to generate queries and the other generates the response. We instruct the user model with… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraChat.texttext-generation100K<n<1M505 likes4.3k downloads3y agoHugging Face12openbmb /VisRAG-Ret-Train-Synthetic-data Dataset Description This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries. Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset. Name Source Description # Pages Textbooks https://openstax.org/ College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.image100K<n<1M20 likes3.3k downloads2y agoHugging Face13openbmb /factnet_factsynset FactSynset Dataset Overview FactSynset is the semantic equivalence layer of FactNet that aggregates similar FactStatements into unified semantic classes with normalized values. It provides a cross-lingual view of semantically equivalent facts, enabling reasoning across language barriers. Paper: https://arxiv.org/abs/2602.03417 Github: https://github.com/yl-shen/factnet Dataset: https://huggingface.co/collections/openbmb/factnet Dataset Format The dataset… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factsynset.tabular1B<n<10B4 likes2.8k downloads5mo agoHugging Face14openbmb /factnet_factsense FactSense Dataset Overview FactSense is the linguistic layer of FactNet that provides multilingual, natural language expressions of facts extracted from Wikipedia pages. Each FactSense instance represents a FactStatement realized in natural text with provenance information. Paper: https://arxiv.org/abs/2602.03417 Github: https://github.com/yl-shen/factnet Dataset: https://huggingface.co/collections/openbmb/factnet Dataset Format The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factsense.tabular1B<n<10B6 likes2.3k downloads5mo agoHugging Face15openbmb /UltraInteract_pair Introduction 📜 Paper 🤗 Eurus Collection 🤗 UltraInteract SFT Preference Learning GitHub Repo UltraInteract is a large-scale, high-quality alignment dataset specifically designed for complex reasoning tasks. For each instruction, it includes a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format (2) multi-turn interaction trajectories with the environment and the critique (3) pairwise data to facilitate preference learning… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraInteract_pair.text100K<n<1M111 likes2k downloads2y agoHugging Face16openbmb /factnet_factstatements FactStatement Dataset Overview FactStatement is the foundational layer of FactNet, a cross-lingual, multi-layered fact knowledge graph. FactStatements are language-neutral, atomic fact units directly mapped from Wikidata statements, forming the core building blocks of the knowledge graph. Paper: https://arxiv.org/abs/2602.03417 Github: https://github.com/yl-shen/factnet Dataset: https://huggingface.co/collections/openbmb/factnet Dataset Format The dataset… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_factstatements.text1B<n<10B7 likes2k downloads5mo agoHugging Face17openbmb /RLAIF-V-Dataset Dataset Card for RLAIF-V-Dataset This dataset was introduced in RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness. GitHub This dataset was also used in MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe News: [2025.09.18] 🎉 Our data is used in the powerful MiniCPM-V 4.5 model, which represents a state-of-the-art end-side MLLM achieving GPT-4o level performance! [2025.03.01] 🎉 RLAIF-V is accepted by CVPR… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLAIF-V-Dataset.imageimage-text-to-text10K<n<100K219 likes1.9k downloads11mo agoHugging Face18openbmb /VisRAG-Ret-Test-ArxivQA Dataset Description This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"])… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-ArxivQA.image1K<n<10K2 likes1.8k downloads2y agoHugging Face19openbmb /VisRAG-Ret-Test-SlideVQA Dataset Description This is a VQA dataset based on Slide Decks from SlideVQA dataset from SlideVQA. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"]) if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-SlideVQA.image1K<n<10K2 likes1.7k downloads2y agoHugging Face20openbmb /VisRAG-Ret-Test-PlotQA Dataset Description This is a VQA dataset based on Scientific Plots from PlotQA dataset from PlotQA. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"]) if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-PlotQA.image10K<n<100K2 likes1.6k downloads2y agoHugging Face21openbmb /VisRAG-Ret-Test-ChartQA Dataset Description This is a VQA dataset based on Charts from ChartQA dataset from ChartQA. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"]) if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-ChartQA.imagen<1K1 likes1.6k downloads2y agoHugging Face22openbmb /VisRAG-Ret-Test-MP-DocVQA Dataset Description This is a VQA dataset based on Industrial Documents from MP-DocVQA dataset from MP-DocVQA. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"]) if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-MP-DocVQA.image1K<n<10K1 likes1.6k downloads2y agoHugging Face23openbmb /VisRAG-Ret-Test-InfoVQA Dataset Description This is a VQA dataset based on Infographics from InfoVQA dataset from InfoVQA. Load the dataset from datasets import load_dataset import csv def load_beir_qrels(qrels_file): qrels = {} with open(qrels_file) as f: tsvreader = csv.DictReader(f, delimiter="\t") for row in tsvreader: qid = row["query-id"] pid = row["corpus-id"] rel = int(row["score"]) if qid in qrels:… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Test-InfoVQA.image1K<n<10K1 likes1.6k downloads2y agoHugging Face24openbmb /UltraInteract_sft Introduction 📜 Paper 🤗 Eurus Collection 🤗 UltraInteract SFT Preference Learning GitHub Repo UltraInteract is a large-scale, high-quality alignment dataset specifically designed for complex reasoning tasks. For each instruction, it includes a preference tree consisting of (1) reasoning chains with diverse planning strategies in a unified format (2) multi-turn interaction trajectories with the environment and the critique (3) pairwise data to facilitate preference learning… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraInteract_sft.text100K<n<1M129 likes1.4k downloads2y agoHugging Face25openbmb /VisRAG-Ret-Train-In-domain-data Dataset Description This dataset is the In-domain part of the training set of VisRAG it includes 122,752 Query-Document (Q-D) Pairs from openly available academic datasets. Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset. Dataset # Q-D Pairs ArXivQA 25,856 ChartQA 4,224 MP-DocVQA 10,624 InfoVQA 17,664 PlotQA 56,192 SlideVQA 8,192 Load the dataset from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-In-domain-data.image100K<n<1M9 likes1.1k downloads2y agoHugging Face26openbmb /FormalVerse MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement FormalVerse is a verified Lean 4 autoformalization dataset released with the paper MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement. Every example is produced by the MathForm pipeline, which retrieves relevant Mathlib knowledge before generation and refines each candidate using Lean compiler diagnostics and… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/FormalVerse.texttext-generation100K<n<1M14 likes762 downloads1mo agoHugging Face27openbmb /MA-ProofBench MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis English | 中文 We introduce MA-ProofBench, to the best of our knowledge, the first formal benchmark for evaluating large language models (LLMs) on theorem proving in Mathematical Analysis. It contains 200 rigorously formalized theorem-proving problems in Lean 4 + Mathlib (v4.28.0), split into two difficulty tiers: Tier Description Source Count Level I Undergraduate… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/MA-ProofBench.texttext-generationn<1K16 likes748 downloads27d agoHugging Face28openbmb /RLHF-V-Dataset Dataset Card for RLHF-V-Dataset Project Page | Paper | GitHub Updates [2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now! [2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now! [2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset.imagetext-generation1K<n<10K73 likes562 downloads2y agoHugging Face29openbmb /RLPR-Train-Dataset Dataset Card for RLPR-Train-Dataset GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and this dataset is accessible at here. Dataset Summary The RLPR-Train-Dataset is a curated collection of 77k high-quality reasoning prompts specifically designed for enhancing Large Language Model (LLM) capabilities in the general domain (non-mathematical). This dataset is derived from the comprehensive collection of prompts from WebInstruct. We… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLPR-Train-Dataset.texttext-generation10K<n<100K30 likes485 downloads1y agoHugging Face30openbmb /InfLLM-V2-data-5B InfLLM-V2 Long-Context Training Dataset with 5B Tokens Project Links: [Paper] [InfLLM-V2 Models] [CUDA Kernel Code] 🚀 About InfLLM-V2 InfLLM-V2 is a native sparse attention framework designed for the efficient processing of long-sequence texts. Its core advantage is the ability to maintain high performance comparable to dense attention in short-text scenarios—without any extra parameters—while seamlessly switching to a sparse mode for long-text scenarios, achieving… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/InfLLM-V2-data-5B.text1M<n<10M36 likes300 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.