CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hltcoe /megawikaMegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser.summarization10M<n<100M42 likes134k downloads2y agoHugging Face02moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes19k downloads1y agoHugging Face03racineai /VDR_MEGA_2 VDR_MEGA_2 Dataset Summary VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.imagequestion-answering1M<n<10M16 likes8.9k downloads10mo agoHugging Face04guodaosun /Mega60k Mega60k: Chart Question Answering Dataset Dataset Overview A multimodal chart question answering dataset featuring charts in multiple formats (CSV, PNG, SVG) and degraded PNG images with components omission, occlusion, blurring, and rotation to enhance robustness evaluation. Languages: English Chart Type Distribution Chart Type Count Chart Type Count Chart Type Count Area 200 Bar 200 Box 200 Bubble 200 Chord 200 Fill-bubble 200 Funnel 200… See the full description on the dataset page: https://huggingface.co/datasets/guodaosun/Mega60k.imagequestion-answering1M<n<10M0 likes4.3k downloads10mo agoHugging Face05OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face06megagonlabs /subjqaSubjQA is a question answering dataset that focuses on subjective questions and answers. The dataset consists of roughly 10,000 questions over reviews from 6 different domains: books, movies, grocery, electronics, TripAdvisor (i.e. hotels), and restaurants.question-answering1K<n<10K16 likes573 downloads3y agoHugging Face07TIGER-Lab /MEGA-Bench MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks [ICLR 2025] 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 🔎 Visualiaztion | 📖 arXiv | GitHub 🔔 News [2025-01]: Paper accepted by ICLR 2025. [2024-10-18]: Initial release of the evaluation code on our Github repo. [2024-10-14]: Paper released on arXiv. ❗❗ Data Information We put the file path of images/videos in HF datasets. Please download the zipped data here. We chose… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MEGA-Bench.imagequestion-answering1K<n<10K23 likes430 downloads1y agoHugging Face08LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes410 downloads8mo agoHugging Face09megagonlabs /holobench HoloBench (Holistic Reasoning Benchmark) HoloBench is a benchmark designed to evaluate the ability of long-context language models (LCLMs) to perform holistic reasoning over extended text contexts. Unlike standard models that retrieve isolated information, HoloBench tests how well LCLMs handle complex reasoning tasks that require aggregating and synthesizing information across multiple documents or large text segments. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/holobench.tabularquestion-answering100K<n<1M3 likes164 downloads2y agoHugging Face10Sidd2229 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/Sidd2229/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes137 downloads7mo agoHugging Face11dhruvR16 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/dhruvR16/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes119 downloads7mo agoHugging Face12nace-ai /MegaWikiQA-v1-multihop MegaWikiQA v1 Multihop Dataset Combined and shuffled Wiki5M-based synthetic multihop QA dataset for hypernetwork / knowledge-injection research. Sources Hop Source dataset Rows 1 nace-ai/wiki5m_1hop_qa_pairs_1M_stratified_with_domain 1,000,000 2 nace-ai/wiki5m_2hop_qa_pairs_noun_v10_domain 1,894,483 3 nace-ai/wiki5m_3hop_qa_pairs_noun_v10_domain 2,393,688 Total shuffled (seed=42) 5,288,171 Schema question, answer hop — 1 / 2 /… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/MegaWikiQA-v1-multihop.textquestion-answering1M<n<10M3 likes89 downloads2mo agoHugging Face13tasal9 /ZamAI-Pashto-Mega-Dataset ZamAI Pashto Mega Dataset Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.texttext-generation1M<n<10M0 likes54 downloads2mo agoHugging Face14introvoyz041 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes51 downloads5mo agoHugging Face15sitrafund /megatrendit-2026 Sitra Megatrendit 2026 Dataset This dataset contains the structured content from Sitra's Megatrends 2026 report (Megatrendit 2026 - Kohti uutta yhteiskuntasopimusta). Why was this dataset created? Sitra's Megatrends report is a key Finnish foresight publication that analyzes global trends shaping society over the next decade. The original report is published as a PDF, which is not an optimal format for LLMs, RAG, and other machine usages. This dataset was created to:… See the full description on the dataset page: https://huggingface.co/datasets/sitrafund/megatrendit-2026.text-generationn<1K2 likes38 downloads8mo agoHugging Face16Aipresso /MEGA-cleaned-prompts Cleaned Prompts Mega Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A comprehensive collection of 2.7 million cleaned English prompts, meticulously processed for training advanced language models and AI systems. 📊 Dataset Statistics Metric Value Total Rows 2,689… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/MEGA-cleaned-prompts.texttext-generation1M<n<10M0 likes36 downloads11mo agoHugging Face17Dragnoz /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/Dragnoz/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M2 likes31 downloads7mo agoHugging Face18celsowm /demons_megaten_fandom_orpo_dpo_englishtextquestion-answering1K<n<10K0 likes28 downloads2y agoHugging Face19CrossNow /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/CrossNow/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M1 likes27 downloads7mo agoHugging Face20tasal9 /ZamAI-Pashto-MegaDataset-v1 ZamAI Pashto Mega Dataset v1 Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-MegaDataset-v1") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-MegaDataset-v1.texttext-generation10K<n<100K0 likes25 downloads2mo agoHugging Face21living-my-best-life /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/living-my-best-life/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes24 downloads7mo agoHugging Face22jumpzakrab /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/jumpzakrab/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes23 downloads6mo agoHugging Face23mk12333 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/mk12333/Medical-Reasoning-SFT-Mega.text-generation1M<n<10M0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.