CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yatin-superintelligence /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M52 likes6.1k downloads6mo agoHugging Face02Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face03Emova-ollm /emova-sft-4m EMOVA-SFT-4M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.imageimage-to-text1M<n<10M6 likes3.1k downloads2y agoHugging Face04zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face05shintaro-ozaki /entity-explanationimagetext-generation100B<n<1T2 likes3k downloads11mo agoHugging Face06ericktwo /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ericktwo/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M1 likes1.4k downloads8mo agoHugging Face07BUAADreamer /llava-en-zh-300kThis dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English Visual Instruction Data from openbmb. You can use it in LLaMA Factory by specifying --dataset llava_150k_en,llava_150k_zh. imagetext-generation100K<n<1M36 likes1.3k downloads2y agoHugging Face08exnihilum /ttrpg-rpg-fandom-com-en ttrpg-rpg-fandom-com-en (RPG Fandom EN Dataset) [Russian version below / Русская версия ниже] Description This dataset contains a complete dump of the English rpg.fandom.com wiki, converted to clean Markdown. It is designed for RAG (Retrieval-Augmented Generation), LLM fine-tuning, and research. Structure markdown/: Cleaned documents with metadata. indexes/documents.jsonl: Global document registry. indexes/chunks.jsonl: Semantic fragments for… See the full description on the dataset page: https://huggingface.co/datasets/exnihilum/ttrpg-rpg-fandom-com-en.imagetext-generationn<1K0 likes1k downloads2mo agoHugging Face09lhpku20010120 /Omni-Edu Omni-Edu — Core V6 SFT mixture 69,999 supervised instruction examples (~158M characters) covering K-12 subject competence, curriculum grounding, diagnostic reasoning, pedagogical action and general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image referenced by the JSONL ships in this repository under images/. This is the system-prompted assembly of the v6 core mixture: every row carries an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.imagetext-generation10K<n<100K1 likes815 downloads6d agoHugging Face10NationalLibraryOfScotland /encyclopaedia-britannica-lance Encyclopaedia Britannica (1771-1860) - Lance Format This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading. Dataset Details Total Pages: 155,388 Total Volumes: 195 Format: Lance (columnar format with blob storage for images) Source: National Library of Scotland (NLS) License: Public Domain (CC0) Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/encyclopaedia-britannica-lance.imageimage-to-text100K<n<1M2 likes783 downloads8mo agoHugging Face11davanstrien /encyclopaedia-britannica-lance-test Encyclopaedia Britannica (1771-1860) - Lance Format This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading. Dataset Details Total Pages: 155,388 Total Volumes: 195 Format: Lance (columnar format with blob storage for images) Source: National Library of Scotland (NLS) License: Public Domain (CC0) Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test.imageimage-to-text100K<n<1M0 likes724 downloads8mo agoHugging Face12etri-vilab /holisafe-benchgated ⚠️ CONTENT WARNING: This dataset contains potentially harmful and sensitive visual content including violence, hate speech, illegal activities, self-harm, sexual content, and other unsafe materials. Images are intended solely for safety research and evaluation purposes. Viewer discretion is strongly advised. HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model (CVPR'26 Findings) 🌐 Website | 📑 Paper 📋 HoliSafe-Bench Dataset… See the full description on the dataset page: https://huggingface.co/datasets/etri-vilab/holisafe-bench.imagevisual-question-answering1K<n<10K11 likes673 downloads4mo agoHugging Face13BlueIsGreen /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/BlueIsGreen/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M11 likes621 downloads6mo agoHugging Face14DEMIRUNC /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes581 downloads6mo agoHugging Face15biglam /europeana_newspapers Dataset Card for Europeana Newspapers Dataset Overview This dataset contains historic newspapers from Europeana, processed and converted to a format more suitable for machine learning and digital humanities research. In total, the collection contains approximately 32 billion tokens across multiple European languages, spanning from the 18th to the early 20th century. Created by the BigLAM initiative, this unofficial version extracts text content from ALTO XML and… See the full description on the dataset page: https://huggingface.co/datasets/biglam/europeana_newspapers.imagetext-generation10M<n<100M59 likes577 downloads12h agoHugging Face16My-Weird-Prompts /episodes My Weird Prompts - Episode Dataset The production record of every episode of the My Weird Prompts podcast: the transcript, links to the published episode, a description of the prompt that started it, and the generation telemetry for how it was made - model, pipeline version, GPU, timings and compute cost. 5,365 episodes. Synced daily from the production database. from datasets import load_dataset ds = load_dataset("My-Weird-Prompts/episodes", split="train") Which… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/episodes.audiotext-generation1K<n<10K1 likes506 downloads19h agoHugging Face17Torenn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/Torenn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M1 likes400 downloads6mo agoHugging Face18ed001 /ds-coder-instruct-v1 Dataset Card for DS Coder Instruct Dataset DS Coder is a dataset for instruction fine tuning of language models. It is a specialized dataset focusing only on data science (eg. plotting, data wrangling, machine learnig models, deep learning, and numerical computations). The dataset contains code examples both in R and Python. The goal of this dataset is to enable creation of small-scale, specialized language model assistants for data science projects. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/ed001/ds-coder-instruct-v1.imagetext-generation10K<n<100K5 likes378 downloads3y agoHugging Face19eduagarcia /cc_news_pt_v2 Dataset Summary This version of the dataset is the portuguese subset from stanford-oval/ccnews. CC-News-PT v2 is a curation of +11 million news articles from CommonCrawl News in the Portuguese language, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and filtered. The process is similar to what HuggingFace's DataTrove does. For license information, please refer to CommonCrawl's Terms of Use.… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/cc_news_pt_v2.imagetext-classification10M<n<100M4 likes375 downloads1y agoHugging Face20gllllll /glsl-opengl-educational-dataset GLSL/OpenGL & WebGPU Universal Educational Dataset for AI Training Curated, statically validated, and educational dataset of GLSL, WGSL, and HLSL shaders, OpenGL/WebGL programs, and real-time graphics pipelines. Key Features Multi-Stage Shaders: Vertex, Fragment, Compute, Geometry, Tessellation. Static Validation: Validated against Khronos glslangValidator. Universal Shading Targets: Multi-target transpilation (WGSL, HLSL, MSL). Rich Annotations: Includes… See the full description on the dataset page: https://huggingface.co/datasets/gllllll/glsl-opengl-educational-dataset.imagetext-generationn<1K1 likes367 downloads1mo agoHugging Face21ppenner /edge-agent-reasoning-websearch-260k Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/ppenner/edge-agent-reasoning-websearch-260k.texttext-generation100K<n<1M0 likes340 downloads4mo agoHugging Face22zt1106 /OpenClaw-EvalMix OpenClaw EvalMix OpenClaw EvalMix is a Harbor-format collection of 360 agent-evaluation tasks across four task families. Each task directory includes task.toml, an instruction, an environment definition, and verifier tests. Composition Family Tasks Local payload clawbench 19 0.00 GiB liveclawbench 134 0.07 GiB pinchbench 147 0.02 GiB wildclawbench 60 14.05 GiB The repository preserves each family at the root so task paths remain direct.… See the full description on the dataset page: https://huggingface.co/datasets/zt1106/OpenClaw-EvalMix.imagequestion-answeringn<1K0 likes336 downloads3mo agoHugging Face23exnihilum /ttrpg-rpg-fandom-com-ru ttrpg-rpg-fandom-com-ru (RPG Fandom RU Dataset) [Russian version below / Русская версия ниже] Описание Этот датасет содержит полный дамп вики rpg.fandom.com/ru/, конвертированный в чистый Markdown. Он предназначен для использования в системах RAG (Retrieval-Augmented Generation), дообучения языковых моделей (LLM) и исследований в области НРИ. Создатель и контакты Вебсайт: exnihilum.info GitHub: exnpub Репозиторий проекта: dataset-rpg-fandom-com-ru… See the full description on the dataset page: https://huggingface.co/datasets/exnihilum/ttrpg-rpg-fandom-com-ru.imagetext-generationn<1K0 likes305 downloads2mo agoHugging Face24JACKYS999 /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/JACKYS999/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes290 downloads4mo agoHugging Face25Bryan35406 /fable-novel-eightsday fable: eightsday 한국어판 제목: 「fable — 여드레날」 License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md. The story of a man who realized the world is one enormous language model. 세상이 하나의 거대한 언어 모델임을 깨달은 남자의 이야기. A complete Korean–English bilingual serialized novel and a section-aligned literary parallel corpus, co-written by a human… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/fable-novel-eightsday.imagetext-generationn<1K1 likes244 downloads1mo agoHugging Face26eQOURSE /jee-advanced-questions JEE Advanced — Question Bank A structured dataset of JEE Advanced examination questions with full worked solutions and diagrams. JEE Advanced questions are more analytical than JEE Main — many are subjective, integer, or numerical-answer type with detailed multi-step solutions. Subsets (PCM): Physics — 50 questions Chemistry — 21 questions Mathematics — 48 questions Structure Organised into subsets by subject and splits (train / test): mathematics/ physics/… See the full description on the dataset page: https://huggingface.co/datasets/eQOURSE/jee-advanced-questions.imagequestion-answeringn<1K0 likes240 downloads3mo agoHugging Face27davanstrien /encyclopaedia-britannica-lance-test2 Encyclopaedia Britannica (1771-1860) - Lance Format This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading. Dataset Details Total Pages: 155,388 Total Volumes: 195 Format: Lance (columnar format with blob storage for images) Source: National Library of Scotland (NLS) License: Public Domain (CC0) Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test2.imageimage-to-text100K<n<1M0 likes234 downloads8mo agoHugging Face28Kirito-Lab /VLM-ExecRouterBench VLM-ExecRouterBench An execution-oriented benchmark for cost-aware open-set VLM routing. Cost-aware routing | Open-set model onboarding | Multimodal, code, and search tasks Overview VLM-ExecRouterBench is an execution-oriented benchmark for routing vision-language model queries to a pool of candidate VLMs. Each sample is executed by multiple candidate models, producing correctness labels, inference costs, metadata… See the full description on the dataset page: https://huggingface.co/datasets/Kirito-Lab/VLM-ExecRouterBench.imagevisual-question-answering10K<n<100K0 likes228 downloads1mo agoHugging Face29EthnicErotic /phenotype-catalog Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.imagetext-classification10K<n<100K1 likes211 downloads3d agoHugging Face30svryn /Edge-Agent-Reasoning-WebSearch-260K Edge Agent Reasoning WebSearch 260K Abstract The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning. Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/svryn/Edge-Agent-Reasoning-WebSearch-260K.texttext-generation100K<n<1M0 likes188 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.