CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hltcoe /megawikaMegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser.summarization10M<n<100M42 likes139k downloads2y agoHugging Face02moondream /megalith-mdqa Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM. imagequestion-answering1M<n<10M28 likes19k downloads1y agoHugging Face03racineai /VDR_MEGA_2 VDR_MEGA_2 Dataset Summary VDR_MEGA_2 is a high-quality multimodal dataset created through the merge of multiple domain-specific datasets with enhanced data processing techniques. This dataset represents our most refined approach to multimodal data generation, incorporating filtering algorithms and improved AI-assisted content generation to deliver superior quality for RAG, DSE, question answering, document search, and vision-language model training tasks.… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_MEGA_2.imagequestion-answering1M<n<10M16 likes8.8k downloads10mo agoHugging Face04guodaosun /Mega60k Mega60k: Chart Question Answering Dataset Dataset Overview A multimodal chart question answering dataset featuring charts in multiple formats (CSV, PNG, SVG) and degraded PNG images with components omission, occlusion, blurring, and rotation to enhance robustness evaluation. Languages: English Chart Type Distribution Chart Type Count Chart Type Count Chart Type Count Area 200 Bar 200 Box 200 Bubble 200 Chord 200 Fill-bubble 200 Funnel 200… See the full description on the dataset page: https://huggingface.co/datasets/guodaosun/Mega60k.imagequestion-answering1M<n<10M0 likes4.3k downloads10mo agoHugging Face05OpenMed /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M100 likes1.8k downloads8mo agoHugging Face06meg-tong /sycophancy-eval SycophancyEval This is a data-only mirror of our GitHub repository https://github.com/meg-tong/sycophancy-eval, which also contains example evaluation code. This repository includes datasets designed to evaluate sycophantic behavior of language models across varied free-form text-generation tasks from our paper Towards Understanding Sycophancy in Language Models. For questions, please email meg at anthropic dot com. BibTeX citation If you would like to cite our work… See the full description on the dataset page: https://huggingface.co/datasets/meg-tong/sycophancy-eval.text-generationn<1K7 likes796 downloads3y agoHugging Face07megagonlabs /subjqaSubjQA is a question answering dataset that focuses on subjective questions and answers. The dataset consists of roughly 10,000 questions over reviews from 6 different domains: books, movies, grocery, electronics, TripAdvisor (i.e. hotels), and restaurants.question-answering1K<n<10K16 likes575 downloads3y agoHugging Face08TIGER-Lab /MEGA-Bench MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks [ICLR 2025] 🌐 Homepage | 🏆 Leaderboard | 🤗 Dataset | 🤗 Paper | 🔎 Visualiaztion | 📖 arXiv | GitHub 🔔 News [2025-01]: Paper accepted by ICLR 2025. [2024-10-18]: Initial release of the evaluation code on our Github repo. [2024-10-14]: Paper released on arXiv. ❗❗ Data Information We put the file path of images/videos in HF datasets. Please download the zipped data here. We chose… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/MEGA-Bench.imagequestion-answering1K<n<10K23 likes428 downloads1y agoHugging Face09LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes428 downloads8mo agoHugging Face10megagonlabs /holobench HoloBench (Holistic Reasoning Benchmark) HoloBench is a benchmark designed to evaluate the ability of long-context language models (LCLMs) to perform holistic reasoning over extended text contexts. Unlike standard models that retrieve isolated information, HoloBench tests how well LCLMs handle complex reasoning tasks that require aggregating and synthesizing information across multiple documents or large text segments. Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/holobench.tabularquestion-answering100K<n<1M3 likes155 downloads2y agoHugging Face11Sidd2229 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/Sidd2229/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes143 downloads7mo agoHugging Face12dhruvR16 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/dhruvR16/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes119 downloads7mo agoHugging Face13nace-ai /MegaWikiQA-v1-multihop MegaWikiQA v1 Multihop Dataset Combined and shuffled Wiki5M-based synthetic multihop QA dataset for hypernetwork / knowledge-injection research. Sources Hop Source dataset Rows 1 nace-ai/wiki5m_1hop_qa_pairs_1M_stratified_with_domain 1,000,000 2 nace-ai/wiki5m_2hop_qa_pairs_noun_v10_domain 1,894,483 3 nace-ai/wiki5m_3hop_qa_pairs_noun_v10_domain 2,393,688 Total shuffled (seed=42) 5,288,171 Schema question, answer hop — 1 / 2 /… See the full description on the dataset page: https://huggingface.co/datasets/nace-ai/MegaWikiQA-v1-multihop.textquestion-answering1M<n<10M3 likes91 downloads2mo agoHugging Face14tasal9 /ZamAI-Pashto-Mega-Dataset ZamAI Pashto Mega Dataset Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.texttext-generation1M<n<10M0 likes52 downloads2mo agoHugging Face15introvoyz041 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes52 downloads5mo agoHugging Face16Aipresso /MEGA-cleaned-prompts Cleaned Prompts Mega Dataset Created by Aipresso LIMITED, London, UK ⚠️ IMPORTANT: By using this dataset, you agree to our Terms of Use You must provide attribution when using this data in publications, research, or commercial products. Dataset Overview A comprehensive collection of 2.7 million cleaned English prompts, meticulously processed for training advanced language models and AI systems. 📊 Dataset Statistics Metric Value Total Rows 2,689… See the full description on the dataset page: https://huggingface.co/datasets/Aipresso/MEGA-cleaned-prompts.texttext-generation1M<n<10M0 likes37 downloads11mo agoHugging Face17sitrafund /megatrendit-2026 Sitra Megatrendit 2026 Dataset This dataset contains the structured content from Sitra's Megatrends 2026 report (Megatrendit 2026 - Kohti uutta yhteiskuntasopimusta). Why was this dataset created? Sitra's Megatrends report is a key Finnish foresight publication that analyzes global trends shaping society over the next decade. The original report is published as a PDF, which is not an optimal format for LLMs, RAG, and other machine usages. This dataset was created to:… See the full description on the dataset page: https://huggingface.co/datasets/sitrafund/megatrendit-2026.text-generationn<1K2 likes37 downloads9mo agoHugging Face18Dragnoz /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/Dragnoz/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M2 likes31 downloads7mo agoHugging Face19celsowm /demons_megaten_fandom_orpo_dpo_englishtextquestion-answering1K<n<10K0 likes27 downloads2y agoHugging Face20CrossNow /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/CrossNow/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M1 likes27 downloads7mo agoHugging Face21tasal9 /ZamAI-Pashto-MegaDataset-v1 ZamAI Pashto Mega Dataset v1 Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 10K<n<100K Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-MegaDataset-v1") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-MegaDataset-v1.texttext-generation10K<n<100K0 likes25 downloads2mo agoHugging Face22jumpzakrab /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/jumpzakrab/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes23 downloads6mo agoHugging Face23living-my-best-life /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/living-my-best-life/Medical-Reasoning-SFT-Mega.texttext-generation1M<n<10M0 likes15 downloads7mo agoHugging Face24mk12333 /Medical-Reasoning-SFT-Mega Medical-Reasoning-SFT-Mega The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning. Dataset Overview Metric Value Total Samples 1,789,998 (after deduplication) Total Tokens ~3.78 Billion Content Tokens ~2.22 Billion Reasoning Tokens ~1.56 Billion Samples with Reasoning 1,789,764 (100.0%) Unique… See the full description on the dataset page: https://huggingface.co/datasets/mk12333/Medical-Reasoning-SFT-Mega.text-generation1M<n<10M0 likes5 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.