CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /MegaMath MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/IFM/MegaMath.texttext-generation100M<n<1B134 likes120k downloads1y agoHugging Face02racineai /VDR_MEGA_MultiDomain_DocRetrieval Visual Document Retrieval Dataset Overview This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks. Dataset Structure The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.imagevisual-document-retrieval1M<n<10M24 likes72k downloads6mo agoHugging Face03moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes19k downloads1y agoHugging Face04drawthingsai /megalith-10mimage1M<n<10M10 likes10k downloads2y agoHugging Face05racineai /VDR_MEGA_2 VDR_MEGA_2 Dataset Summary VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.imagequestion-answering1M<n<10M16 likes8.8k downloads10mo agoHugging Face06NP235 /MegaMath MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.texttext-generation100M<n<1B0 likes8.4k downloads3mo agoHugging Face07ShinMK3 /Mega-Brain-Distill Mega-Brain-Distill Curated merge of the top 10% highest-scoring examples from 584 community-uploaded LLM distillation/reasoning-trace datasets on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces, etc.), deduplicated within and across all of them — many of these source repos are the same underlying dump re-uploaded by different users. Auto-generated by run.py — do not hand-edit, it will be overwritten on the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.tabulartext-generation10K<n<100K2 likes5.4k downloads2mo agoHugging Face08tencent /MegaStyle-1.4MDataset of MegaStyle and MegaStyle++. MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles. MegaStyle++-8M further scales up the style space through a hierarchical style definition. It covers 150K overall style… See the full description on the dataset page: https://huggingface.co/datasets/tencent/MegaStyle-1.4M.imagetext-to-image1M<n<10M54 likes4.6k downloads22d agoHugging Face09MegaScience /MegaScience MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Code: https://github.com/GAIR-NLP/MegaScience Project Page: https://huggingface.co/MegaScience MegaScience is a large-scale mixture of high-quality open-source datasets consisting of 1.25 million instances. We first collect multiple public datasets, then conduct comprehensive ablation studies across different data selection methods to identify the optimal approach for each dataset, thereby… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/MegaScience.texttext-generation1M<n<10M134 likes4k downloads1y agoHugging Face10meganwei /syntheory Dataset Card for SynTheory Dataset Summary SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions). Each of these 7 concepts has its own config. tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time). time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.audioaudio-classification100K<n<1M13 likes3.4k downloads2y agoHugging Face11OctoThinker /MegaMath-Web-Pro-Max OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year; Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data; Step 3: Training a fasttext carefully with proper preprocessing; Step 4: Filtering documents with a threshold (i.e., 0.4); Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.tabular10M<n<100M41 likes2.7k downloads1y agoHugging Face12MegaScience /TextbookReasoning MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Dataset Description Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.texttext-generation100K<n<1M33 likes1.9k downloads1y agoHugging Face13OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face14celestialformeventually /megalith-10mimage1M<n<10M0 likes950 downloads23d agoHugging Face15madebyollin /megalith-10m 🗿 Megalith-10m What is Megalith-10m? Megalith-10m is a dataset of ~10 million links to Flickr images that were categorized as "photo" with license info of: No known copyright restrictions (Flickr commons), or United States Government Work, or Public Domain Dedication (CC0), or Public Domain Mark What's the intended use of Megalith-10m? Megalith-10m is intended to contain only links to wholesome unedited uncopyrighted photographs - the sort of… See the full description on the dataset page: https://huggingface.co/datasets/madebyollin/megalith-10m.image1M<n<10M108 likes864 downloads4mo agoHugging Face16di-zhang-fdu /MegaTerminal MegaTerminal MegaTerminal is a Harbor-style terminal task-folder dataset assembled from the terminal task sources collected for task matching experiments. The task folders are packed into parquet/ as a custom blob layout (megatask-folder-parquet-v1); each row stores the files of one task folder as binary blobs. Restore the on-disk tasks/<task-id>/ tree with python scripts/deparquetize_megaterminal.py --parquet-dir parquet --out-dir MegaTerminal-restored. Each task lives under… See the full description on the dataset page: https://huggingface.co/datasets/di-zhang-fdu/MegaTerminal.tabular10K<n<100K0 likes577 downloads3mo agoHugging Face17RosettaCommons /MegaScale Mega-scale experimental analysis of protein folding stability in biology and design The full MegaScale dataset contains 1,841,285 thermodynamic folding stability measurements using cDNA display proteolysis of natural and designed proteins. From these 776,298 high-quality folding stabilities (dataset2) cover all single amino acid variants and selected double mutants of 331 natural and 148 de novo designed protein domains 40–72 amino acids in length. Of these mutations, 607,839 have… See the full description on the dataset page: https://huggingface.co/datasets/RosettaCommons/MegaScale.tabular1M<n<10M5 likes480 downloads2y agoHugging Face18TIGER-Lab /MEGA-Bench MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks [ICLR 2025] 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 🔎 Visualiaztion | 📖 arXiv | GitHub 🔔 News [2025-01]: Paper accepted by ICLR 2025. [2024-10-18]: Initial release of the evaluation code on our Github repo. [2024-10-14]: Paper released on arXiv. ❗❗ Data Information We put the file path of images/videos in HF datasets. Please download the zipped data here. We chose… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MEGA-Bench.imagequestion-answering1K<n<10K23 likes428 downloads1y agoHugging Face19LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes428 downloads8mo agoHugging Face20sawradip /bn-asr-mega-open-dataaudio1M<n<10M3 likes421 downloads3y agoHugging Face21YouAreSpecialToMe /filtered_MegaMathtext10M<n<100M0 likes421 downloads1y agoHugging Face22lehduong /megamath-synthtext10M<n<100M1 likes415 downloads1y agoHugging Face23ASSERT-KTH /megadiff Megadiff, a dataset of source code changes If you use Megadiff, please cite the following technical report: "Megadiff: A Dataset of 600k Java Source Code Changes Categorized by Diff Size". Technical Report 2108.04631, Arxiv; 2021. @techreport{megadiff, TITLE = {{Megadiff: A Dataset of 600k Java Source Code Changes Categorized by Diff Size}}, AUTHOR = {Martin Monperrus and Matias Martinez and He Ye and Fernanda Madeiral and Thomas Durieux and Zhongxing Yu}, URL =… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/megadiff.text100K<n<1M2 likes387 downloads3y agoHugging Face24MohamedGomaa30 /MasriAudio-Mega-v0audio100K<n<1M2 likes365 downloads8mo agoHugging Face25lehduong /megamath-web-protabular10M<n<100M0 likes329 downloads1y agoHugging Face26meganariley /daily-aqi Daily AQI — US Air Quality Sharing datasets helps agents analyze them — giving everyone the ability to make sense of complex data. Official US EPA air quality data updated daily. All 6 NAAQS criteria pollutants, hourly granularity, ~1,500 monitoring stations across the United States. Explore it interactively in the Daily AQI Space. Dataset Structure readings.parquet — hourly readings One row per monitoring station per hour per pollutant. Column… See the full description on the dataset page: https://huggingface.co/datasets/meganariley/daily-aqi.tabular100M<n<1B0 likes327 downloads5mo agoHugging Face27komats /mega-ssum Mega-SSum A large-scale English sentence-wise speech summarization (Sen-SSum) dataset Consists of 3.8M+ synthesized speech, transcription, summary triplets Derived from the Gigaword dataset Rush+2015 Overview The dataset is divided into five splits: train/core/dev/eval/duc2003. (See below table) We added a new evaluation split "test" for in-domain evaluation. The train split is here: MegaSSum(train). orig. data split #samples #speakers total dur. (hrs) ave.… See the full description on the dataset page: https://huggingface.co/datasets/komats/mega-ssum.audio10K<n<100K3 likes314 downloads2y agoHugging Face28jwu323 /MegaTerminal-Hard-1K MegaTerminal-Hard-1K The 1,000 highest-quality, hardest terminal-agent tasks selected from the 11,599 tasks in di-zhang-fdu/MegaTerminal. Every task is a complete, self-contained Harbor / Terminal-Bench task folder — instruction, container definition, verifier, and reference solution — restorable byte-for-byte from the Parquet blobs in parquet/. Tasks 1,000 (top 9.8% of the deduplicated upstream pool) Restored size 156 MB across 10,119 files With reference… See the full description on the dataset page: https://huggingface.co/datasets/jwu323/MegaTerminal-Hard-1K.tabularother1K<n<10K0 likes312 downloads2mo agoHugging Face29nfsrulesFR /mega-moledit-large MEGA: A Large-Scale Molecular Editing Dataset for Guided-Action Optimization Large-scale annotated molecular editing dataset with 57M examplesfor training models to modify molecular structures based on natural language instructions. Paper: MEGA: A Large-Scale Molecular Editing Dataset for Guided-Action Optimization Official Repository: https://github.com/nfsrules/MEGA-moledit Dataset Structure Each example will contain: task_id: Task identifier prompt: Natural… See the full description on the dataset page: https://huggingface.co/datasets/nfsrulesFR/mega-moledit-large.tabulartext-generation10M<n<100M0 likes302 downloads10mo agoHugging Face30Spawning /megalith-cc0 Megalith-CC0 A CC0-filtered version of the Megalith-10m dataset. The images have also been persisted to an independent public S3 bucket, supported by the AWS Open Data Registry program, for durability. Why filter by CC0? The images in Megalith-10m, having been gathered from Flickr, have attached licenses of CC0 and public domain. However, it is not clear if users assigning the public domain license to their works understand the implications of the public domain… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/megalith-cc0.image1M<n<10M3 likes300 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.