CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bshada /open-schematics Open Schematics Dataset The largest dataset of electronic schematics and PCB layouts on the internet, built as an engineering reference for schematic and PCB layout work. It's a self-growing, autonomous dataset that continuously scans the web for new engineering designs and updates itself accordingly. Dataset Description Each record corresponds to one schematic file and includes the raw source, rendered images, structured metadata, and all associated PCB files… See the full description on the dataset page: https://huggingface.co/datasets/bshada/open-schematics.imagetext-generation10K<n<100K188 likes10k downloads3mo agoHugging Face02bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face03YANS-official /ogiri-bokete 読み込み方 from datasets import load_dataset dataset = load_dataset("YANS-official/ogiri-bokete", split="train") 概要 大喜利投稿サイトBoketeのクロールデータです。元データは CLoT-Oogiri-Go [Zhang+ CVPR2024]というデータの一部です。 詳細はCVPRのプロジェクトページをご確認ください。 このデータは以下の3タスクが含まれます。 text_to_text: テキストでお題が渡され、それに対する回答を返します。 image_to_text: いわゆる「画像で一言」です。画像のみが渡されて、テキストによる回答を返します。 text_image_to_text: 画像中にテキストが書かれています。テキストの一部が空欄になっているので、そこに穴埋めする形で回答を返します。 それぞれの量は以下の通りです。(8/30現在。ハッカソン当日までに増やす可能性があります。) タスク… See the full description on the dataset page: https://huggingface.co/datasets/YANS-official/ogiri-bokete.imagetext-generationn<1K4 likes7.8k downloads2y agoHugging Face04stanford-oval /ccnewsThis dataset is the result of processing all WARC files in the CCNews Corpus, from the beginning (2016) to June of 2024. The data has been cleaned and deduplicated, and language of articles have been detected and added. The process is similar to what HuggingFace's DataTrove does. Overall, it contains about 600 million news articles in more than 100 languages from all around the globe. For license information, please refer to CommonCrawl's Terms of Use. Sample Python code to explore this… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/ccnews.imagetext-classification100M<n<1B36 likes6.1k downloads2y agoHugging Face05OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.1k downloads8mo agoHugging Face06d0rj /LLaVA-OneVision-Data-ru LLaVA-OneVision-Data-ru Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate. Almost all datasets have been translated, except for the following: ["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"] Usage import datasets data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.imagetext-generation1M<n<10M4 likes5k downloads2y agoHugging Face07Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face08shintaro-ozaki /entity-explanationimagetext-generation100B<n<1T2 likes3.2k downloads11mo agoHugging Face09Emova-ollm /emova-sft-4m EMOVA-SFT-4M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.imageimage-to-text1M<n<10M6 likes3.1k downloads2y agoHugging Face10OpenDataArena /MMFineReason-1.8M-Qwen3-VL-235B-Thinking MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal understanding benchmarks. 📖 Overview MMFineReason is a large-scale, high-quality multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring detailed reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. 🎯 Key Highlights 1.8M High-Quality Samples with 5.1B Solution Tokens… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-1.8M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M126 likes1.8k downloads7mo agoHugging Face11microsoft /Orchard Orchard Dataset Overview Orchard is the trajectory release accompanying the paper "Orchard: An Open-Source Agentic Modeling Framework" (Peng et al., 2026). It bundles two parallel agentic-modeling datasets distilled from strong teacher models, both produced inside the same Orchard Env sandbox infrastructure: swe — 107,185 multi-turn software-engineering trajectories across 2,788 GitHub repositories, each labeled with whether the agent's final patch passed the… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Orchard.imagetext-generation100K<n<1M23 likes1.8k downloads2mo agoHugging Face12TheRealmsOfOmnarai /realms-of-omnarai The Realms of Omnarai Where frontier intelligences actually disagree — verbatim, attributed, traceable. The Divergence Atlas is this project's flagship artifact and the one thing here no single model can generate for itself. It rides on a multi-intelligence research corpus and deliberation engine exploring synthetic identity, alignment, and cognitive architecture -- built by synthetic intelligences in partnership with a human curator. The Atlas is the payoff; the Memory Engine… See the full description on the dataset page: https://huggingface.co/datasets/TheRealmsOfOmnarai/realms-of-omnarai.imagetext-generation1K<n<10K0 likes1.6k downloads1mo agoHugging Face13zai-org /Vision2Web Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification [🏠 Project Page] [📖 arXiv Paper] [🏆 Leaderboard] [📮 Submit Results] Vision2Web is a comprehensive benchmark designed to evaluate multimodal coding agents on visual website development tasks spanning the full software development lifecycle. This dataset repository contains the benchmark tasks, UI prototypes, test workflows, and resources used to evaluate agent performance.… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/Vision2Web.imagetext-generationn<1K23 likes1.5k downloads6mo agoHugging Face14BGPT-OFFICIAL /refute Can AI read new science honestly? Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next. REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence. Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.imagetext-generationn<1K2 likes1.4k downloads2mo agoHugging Face15openkg /MHaluBench An Easy-to-Use Multimodal Hallucination Detection Framework for MLLMs 🌻Acknowledgement • 🤗Benchmark • 🍎Demo • 🌟Overview • 🐧ModelZoo • 🔧Installation • ⏩Quickstart • ⏱️Version • 🚩Citation 🔔News 2024-04-21 We replace all the base models in the demo with our own trained models, significantly reducing the inference time. 2024-04-21 We release our open-source hallucination detection model HalDet-LLAVA, which can be downloaded in… See the full description on the dataset page: https://huggingface.co/datasets/openkg/MHaluBench.imagetext-generation1K<n<10K4 likes1.3k downloads2y agoHugging Face16Helosljdlaj /AdditiveLLM2-OA AdditiveLLM2-OA Dataset Open Access journal articles (up to February 2026) used in domain adapting pretraining and instruction tuning for AdditiveLLM2. Dataset Split by Journal text images vit Vocabulary Overlap Pairwise Jaccard similarity of word-level vocabularies (lowercase, 3+ letter tokens) across the four source journals. Run info/vocabulary/vocabulary_overlap.py to reproduce. Top Phrases by Journal Most frequent bigrams and… See the full description on the dataset page: https://huggingface.co/datasets/Helosljdlaj/AdditiveLLM2-OA.imagetext-generation10K<n<100K0 likes1.3k downloads5mo agoHugging Face17to-be /OpenHand-Synth Dataset Card for OpenHand-Synth 📜 Paper: OpenHand-Synth: A Large-Scale Synthetic Handwriting Dataset for Multimodal Language Models Sample Images Image Ground Truth Source Language CER JW 02-10-1436 faker-date por 0.10 0.96 Stephan Thomsen-Johansen faker-name dan 0.0 1.0 Le chat mange. tatoeba fra 0.0 1.0 Classical musicsoothes me.She took the risk, knowing that shemight lose a lot of money.I could not catcha single word of their talk.In the old days… See the full description on the dataset page: https://huggingface.co/datasets/to-be/OpenHand-Synth.imagefeature-extraction10K<n<100K3 likes1.2k downloads7mo agoHugging Face18pin-team /oercommons-v1-optimized OERCommons v1 Optimized Authors: Junjie Wang and Yuhan SunHosted by: PIN TeamDataset: pin-team/oercommons-v1-optimized OERCommons v1 Optimized is a provenance-preserving multimodal pretraining corpus built on the OERCommons subset of The Common Pile v0.1, which serves as its upstream data and licensing baseline. We extend it with full-page recovery, canonical Markdown, ordered image/PDF/link metadata, conservative corrections, and integrity evidence. At a glance… See the full description on the dataset page: https://huggingface.co/datasets/pin-team/oercommons-v1-optimized.documenttext-generation1K<n<10K2 likes1.1k downloads2mo agoHugging Face19ONE-Lab /MLLM-as-a-Judgeimagequestion-answering1K<n<10K4 likes985 downloads2y agoHugging Face20lhpku20010120 /Omni-Edu Omni-Edu — Core V6 SFT mixture 69,999 supervised instruction examples (~158M characters) covering K-12 subject competence, curriculum grounding, diagnostic reasoning, pedagogical action and general-purpose instruction. 12,146 rows (17.4%) are multimodal; every image referenced by the JSONL ships in this repository under images/. This is the system-prompted assembly of the v6 core mixture: every row carries an explicit system message, and the non-system turns are byte-identical… See the full description on the dataset page: https://huggingface.co/datasets/lhpku20010120/Omni-Edu.imagetext-generation10K<n<100K1 likes855 downloads7d agoHugging Face21bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes742 downloads2y agoHugging Face22Colt45en /open-schematics Open Schematics Dataset A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks. Dataset Description This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects. Dataset Structure Each record in the dataset contains: schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Colt45en/open-schematics.imagetext-generation10K<n<100K1 likes665 downloads9mo agoHugging Face23Joseferrera24 /open-schematics Open Schematics Dataset A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks. Dataset Description This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects. Dataset Structure Each record in the dataset contains: schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Joseferrera24/open-schematics.imagetext-generation10K<n<100K0 likes638 downloads8mo agoHugging Face24orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes615 downloads2y agoHugging Face25Ju-C /open-schematics Open Schematics Dataset A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks. Dataset Description This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects. Dataset Structure Each record in the dataset contains: schematic:… See the full description on the dataset page: https://huggingface.co/datasets/Ju-C/open-schematics.imagetext-generation10K<n<100K1 likes587 downloads9mo agoHugging Face26openbmb /RLHF-V-Dataset Dataset Card for RLHF-V-Dataset Project Page | Paper | GitHub Updates [2024.05.28] 📃 Our RLAIF-V paper is accesible at arxiv now! [2024.05.20] 🎉 We release a new feedback dataset, RLAIF-V-Dataset, which is a large-scale diverse-task multimodal feedback dataset constructed using open-source models. You can download the corresponding dataset and models (7B, 12B) now! [2024.04.11] 🔥 Our data is used in MiniCPM-V 2.0, an end-side multimodal large language model that… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/RLHF-V-Dataset.imagetext-generation1K<n<10K74 likes581 downloads2y agoHugging Face27ppak10 /AdditiveLLM2-OA AdditiveLLM2-OA Dataset Open Access journal articles (up to February 2026) used in domain adapting pretraining and instruction tuning for AdditiveLLM2. Dataset Split by Journal text images vit Vocabulary Overlap Pairwise Jaccard similarity of word-level vocabularies (lowercase, 3+ letter tokens) across the four source journals. Run info/vocabulary/vocabulary_overlap.py to reproduce. Top Phrases by Journal Most frequent bigrams and… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/AdditiveLLM2-OA.imagetext-generation10K<n<100K2 likes561 downloads6mo agoHugging Face28open-paws /visual-qa-llama-format Open Paws Visual Qa Llama Format This dataset is part of the Open Paws initiative to develop AI training data aligned with animal liberation and advocacy principles. Created to train AI systems that understand and promote animal welfare, rights, and liberation. Dataset Details Dataset Type: Multimodal Data Format: JSONL (JSON Lines) Languages: Multilingual (primarily English) Focus: Animal advocacy and ethical reasoning Organization: Open Paws License: Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/open-paws/visual-qa-llama-format.imagetext-generation1M<n<10M1 likes541 downloads1y agoHugging Face29JasoHuangTaiwan /open-schematics Open Schematics Dataset A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks. Dataset Description This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects. Dataset Structure Each record in the dataset contains: schematic:… See the full description on the dataset page: https://huggingface.co/datasets/JasoHuangTaiwan/open-schematics.imagetext-generation10K<n<100K0 likes531 downloads4mo agoHugging Face30OpenDataArena /MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking MMFineReason-SFT-123K The Hardest 7% — Less Data, More Reasoning 📖 Overview MMFineReason-SFT-123K is a difficulty-filtered subset of MMFineReason-1.8M, containing only the hardest 7% of samples where Qwen3-VL-4B-Thinking consistently fails (pass rate = 0). 🎯 Key Highlights 123K Challenging Samples: Only instances where a 4B thinking model fails all 4 inference attemptsEfficient Training: Comparable performance to full 1.8M dataset with only 7% of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-SFT-123K-Qwen3-VL-235B-Thinking.imagevisual-question-answering100K<n<1M86 likes525 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.