CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /MegaMath MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/IFM/MegaMath.texttext-generation100M<n<1B134 likes141k downloads1y agoHugging Face02hltcoe /megawikaMegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser.summarization10M<n<100M42 likes134k downloads2y agoHugging Face0386Cao /MegaPairs-Standard MegaPairs-Standard (Standardized Version) Dataset Summary This is a standardized, high-efficiency version of the JUNJIE99/MegaPairs dataset. Why use this version? The original dataset is distributed as a massive Tar archive containing millions of images, accompanied by a separate JSONL annotation file. The Problem: Using the original format requires extracting terabytes of small files (which can exhaust disk inodes) or writing complex logic to read from archives. It… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/MegaPairs-Standard.imageimage-to-text10M<n<100M1 likes105k downloads10mo agoHugging Face04racineai /VDR_MEGA_MultiDomain_DocRetrieval Visual Document Retrieval Dataset Overview This dataset is designed for training visual document retrieval models. It combines multiple datasets from the VDR series, Colpali, and LlamaIndex to create the most comprehensive training resource for visual document retrieval tasks. Dataset Structure The dataset contains structured fields including unique identifiers with string lengths ranging from 45 to 50 characters, search query text with variable lengths between… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_MultiDomain_DocRetrieval.imagevisual-document-retrieval1M<n<10M24 likes72k downloads6mo agoHugging Face05hwjiang /MegaSynth12 likes23k downloads1y agoHugging Face06moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes21k downloads1y agoHugging Face07bop-benchmark /megaposetext1M<n<10M1 likes11k downloads2y agoHugging Face08vctvct123 /Megadepthimage100K<n<1M2 likes11k downloads6mo agoHugging Face09drawthingsai /megalith-10mimage1M<n<10M10 likes11k downloads2y agoHugging Face10racineai /VDR_MEGA_2 VDR_MEGA_2 Dataset Summary VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.imagequestion-answering1M<n<10M16 likes8.9k downloads10mo agoHugging Face11NP235 /MegaMath MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math-focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data: We re-extracted mathematical documents from Common Crawl with math-oriented HTML optimizations, fasttext-based filtering and deduplication, all for acquiring higher-quality data on… See the full description on the dataset page: https://huggingface.co/datasets/NP235/MegaMath.texttext-generation100M<n<1B0 likes8.5k downloads2mo agoHugging Face12jhu-clsp /megawika-2 MegaWika 2 MegaWika 2 is an improved multilingual text dataset containing a structured view of Wikipedia articles, the web sources they cite, source text quality estimates, article text translations, and additional article enrichments. Note: Web citations (sources) in the HuggingFace dataset do not include scraped source text; use rehydrate-citations.py to rehydrate them. The initial data release is based on Wikipedia dumps from May 1, 2024. In total, the data contains about 77… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/megawika-2.10M<n<100M6 likes7k downloads7mo agoHugging Face13Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K17 likes6.7k downloads10h agoHugging Face14guodaosun /Mega60k Mega60k: Chart Question Answering Dataset Dataset Overview A multimodal chart question answering dataset featuring charts in multiple formats (CSV, PNG, SVG) and degraded PNG images with components omission, occlusion, blurring, and rotation to enhance robustness evaluation. Languages: English Chart Type Distribution Chart Type Count Chart Type Count Chart Type Count Area 200 Bar 200 Box 200 Bubble 200 Chord 200 Fill-bubble 200 Funnel 200… See the full description on the dataset page: https://huggingface.co/datasets/guodaosun/Mega60k.imagequestion-answering1M<n<10M0 likes5.1k downloads10mo agoHugging Face15ShinMK3 /Mega-Brain-Distill Mega-Brain-Distill Curated merge of the top 10% highest-scoring examples from 584 community-uploaded LLM distillation/reasoning-trace datasets on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces, etc.), deduplicated within and across all of them — many of these source repos are the same underlying dump re-uploaded by different users. Auto-generated by run.py — do not hand-edit, it will be overwritten on the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.tabulartext-generation10K<n<100K2 likes5k downloads2mo agoHugging Face16OctoThinker /MegaMath-Web-Pro-Max OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample millions of documents from the MegaMath-Web corpus, stratified by publication year; Step 2: Annotate them using Llama-3.1-70B-instruct with a scoring prompt from FineMath and prepare the seed data; Step 3: Training a fasttext carefully with proper preprocessing; Step 4: Filtering documents with a threshold (i.e., 0.4); Step 5:… See the full description on the dataset page: https://huggingface.co/datasets/OctoThinker/MegaMath-Web-Pro-Max.tabular10M<n<100M41 likes4.3k downloads1y agoHugging Face17tencent /MegaStyle-1.4MDataset of MegaStyle and MegaStyle++. MegaStyle-1.4M is a large-scale style dataset built through a scalable pipeline that leverages consistent text-to-image style mapping of Qwen-Image. It combines 170K curated style prompts with 400K content prompts to generate 1.4M high-quality images that share strong intra-style consistency while covering diverse fine-grained styles. MegaStyle++-8M further scales up the style space through a hierarchical style definition. It covers 150K overall style… See the full description on the dataset page: https://huggingface.co/datasets/tencent/MegaStyle-1.4M.imagetext-to-image1M<n<10M54 likes4.2k downloads20d agoHugging Face18MegaScience /MegaScience MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Code: https://github.com/GAIR-NLP/MegaScience Project Page: https://huggingface.co/MegaScience MegaScience is a large-scale mixture of high-quality open-source datasets consisting of 1.25 million instances. We first collect multiple public datasets, then conduct comprehensive ablation studies across different data selection methods to identify the optimal approach for each dataset, thereby… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/MegaScience.texttext-generation1M<n<10M134 likes4k downloads1y agoHugging Face19xiaoxiaxu /highresolution-laioncoco-aesthetic-MEGThis dataset is filtered from laioncoco-aesthetic, which is used for academic research on mobile edge generation (MEG). It includes high-resolution 1024-by-1024 text-to-image samples generated by a distilled SDXL with 4-12 denoising steps. The dataset mainly involves the following fields: caption: The text prompt of the image. image: The target image corresponding to the prompt. diffusion: The generative results of the distilled SDXL. latents: The latent features of the distilled SDXL. tabulartext-to-image1K<n<10K0 likes3.9k downloads2y agoHugging Face20cheryyunl /3D-data-megatron0 likes3.7k downloads1y agoHugging Face21meganwei /syntheory Dataset Card for SynTheory Dataset Summary SynTheory is a synthetic dataset of music theory concepts, specifically rhythmic (tempos and time signatures) and tonal (notes, intervals, scales, chords, and chord progressions). Each of these 7 concepts has its own config. tempos consist of 161 total integer tempos (bpm) ranging from 50 BPM to 210 BPM (inclusive), 5 percussive instrument types (click_config_name), and 5 random start time offsets (offset_time). time_signatures… See the full description on the dataset page: https://huggingface.co/datasets/meganwei/syntheory.audioaudio-classification100K<n<1M13 likes3.1k downloads2y agoHugging Face22datablations /python-megatron1 likes3k downloads3y agoHugging Face23amphora /hephaestus-ccx-runs-megarepo1 likes2.5k downloads2mo agoHugging Face24MeghanaKap /miomio_cp1_cache0 likes2.3k downloads2mo agoHugging Face25Haitao999 /things-meg THINGS-MEG This dataset is a processed version of THINGS-MEG, derived from the paper Bridging the Vision-Brain Gap with an Uncertainty-Aware Blur Prior (CVPR 2025). In this version, the MEG data is stored in float16 format, reducing the storage size by half. The original official dataset can be accessed from the OSF repository. Original official dataset: THINGS-data, a multimodal collection of large-scale datasets for investigating object representations in human brain and… See the full description on the dataset page: https://huggingface.co/datasets/Haitao999/things-meg.image10K<n<100K0 likes1.9k downloads6mo agoHugging Face26MegaScience /TextbookReasoning MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning Dataset Description Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while neglecting the scientific domain, largely due to the absence of open, large-scale, high-quality, verifiable scientific reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MegaScience/TextbookReasoning.texttext-generation100K<n<1M33 likes1.8k downloads1y agoHugging Face27OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.7k downloads8mo agoHugging Face28lsxi77777 /MegaDepth-Syn MegaDepth-Syn Dataset The MegaDepth-Syn Dataset is generated from the MegaDepth dataset using our MINIMA data engine, which contains for extra 6 modalities: infrared, depth, event, normal, sketch, and paint. Abstract Image matching for both cross-view and cross-modality plays a critical role in multimodal perception. In practice, the modality gap caused by different imaging systems/styles poses great challenges to the matching task. Existing works try to extract… See the full description on the dataset page: https://huggingface.co/datasets/lsxi77777/MegaDepth-Syn.imageimage-feature-extraction1K<n<10K3 likes1.4k downloads1y agoHugging Face29BangumiBase /megaminocafeterrace Bangumi Image Base of Megami No Café Terrace This is the image base of bangumi Megami no Café Terrace, we detected 80 characters, 8688 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megaminocafeterrace.1K<n<10K0 likes1.2k downloads2y agoHugging Face30BangumiBase /megatonkyuumusashi Bangumi Image Base of Megaton-kyuu Musashi This is the image base of bangumi Megaton-kyuu Musashi, we detected 81 characters, 5660 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/megatonkyuumusashi.image1K<n<10K0 likes1.2k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.