CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes135k downloads1y agoHugging Face02bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face03BrainAlign /brain-lm-alignment-ds002236 Brain–language-model alignment: ds002236 (whole-brain) Lytle et al. 2020 — orthographic, phonological and semantic word processing in school-aged children (8.7–15.5), auditory and visual. Paper: https://pubmed.ncbi.nlm.nih.gov/31956678/ Data: https://openneuro.org/datasets/ds002236/versions/1.0.1 Generated: 2026-09-24 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number in… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds002236.documentn<1K0 likes4.6k downloads35m agoHugging Face04BrainAlign /brain-lm-alignment-ds006239 Brain–language-model alignment: ds006239 (whole-brain) Wang et al. 2025 — word-level phonological and semantic reading tasks in children and adolescents aged 10–17. Paper: https://www.sciencedirect.com/science/article/pii/S2352340925009692 Data: https://openneuro.org/datasets/ds006239/versions/1.0.5 Generated: 2026-09-24 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds006239.documentn<1K2 likes4.2k downloads11m agoHugging Face05Linzhan /UniML3D UniML3D UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.imagetext-to-3d10K<n<100K8 likes3.8k downloads11h agoHugging Face06HabibaAbderrahim /Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-DatasetTunisian Proverbs with Image Associations: A Cultural and Linguistic Dataset Description This dataset explores the rich oral tradition of Tunisian proverbs mapped into text format, pairing each with contextual explanations, English translations both word-to-word and it's equivalent Target Language dynamic, Automated prompt and AI-generated visual interpretations. It bridges linguistic, cultural, and visual modalities making it valuable for tasks in cross-cultural NLP, generative… See the full description on the dataset page: https://huggingface.co/datasets/HabibaAbderrahim/Tunisian-Proverbs-with-Image-Associations-A-Cultural-and-Linguistic-Dataset.imagetranslationn<1K0 likes2.9k downloads1y agoHugging Face07BrainAlign /brain-lm-alignment-ds001894 Brain–language-model alignment: ds001894 (whole-brain) Lytle et al. 2019 — longitudinal word-level phonological processing in children scanned twice, at roughly 10 and 12 years old. Paper: https://www.nature.com/articles/s41597-019-0338-5 Data: https://openneuro.org/datasets/ds001894/versions/1.4.2 Generated: 2026-09-24 Pipeline: https://github.com/suchirsalhan/cdl-representations-brains-babylms Read this first: does the measurement work? Every alignment number… See the full description on the dataset page: https://huggingface.co/datasets/BrainAlign/brain-lm-alignment-ds001894.documentn<1K0 likes2.4k downloads37m agoHugging Face08leibnitz-lab /mdsaimagen<1K0 likes1.6k downloads1y agoHugging Face09linxy97 /genhome3d-1280 GenHome3D-1280 1,280 validated household and spatial-design assets in USDZ format, organized across 64 categories. Explore the visual catalog · Browse the GitHub repository · Download the versioned release · Read the generation method Dataset summary Assets 1,280 Categories 64 Assets per category 20 Runtime format USDZ Units Meters Asset license CC BY 4.0 Technical validation 1,280/1,280 pass Package validation 1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.3d1K<n<10K1 likes1.4k downloads2mo agoHugging Face10leodriesch /open-hdri-1k Open HDRI 1K A consolidated, public-domain (CC0-1.0) collection of 3,491 equirectangular HDR environment maps at 1K resolution, gathered from five free HDRI libraries: Poly Haven, BlenderKit, ambientCG, CGEES and Open HDRI. Every map is stored as a linear, high-dynamic-range .exr file alongside a tonemapped .jpg preview, with a per-asset metadata row (dimensions, source, author, license, SHA-256 checksum and tags). Contents Source Assets Author(s) License… See the full description on the dataset page: https://huggingface.co/datasets/leodriesch/open-hdri-1k.imageimage-to-image1K<n<10K1 likes1.4k downloads2mo agoHugging Face11liupf /ChEBI-20-MM ChEBI-20-MM Dataset Overview The ChEBI-20-MM is an extensive and multi-modal benchmark developed from the ChEBI-20 dataset. It is designed to provide a comprehensive benchmark for evaluating various models' capabilities in the field of molecular science. This benchmark integrates multi-modal data, including InChI, IUPAC, SELFIES, and images, making it a versatile tool for a wide range of molecular tasks. Dataset Description ChEBI-20-MM is an expansion of the… See the full description on the dataset page: https://huggingface.co/datasets/liupf/ChEBI-20-MM.imagetext-generation10K<n<100K9 likes1.1k downloads2y agoHugging Face12bonadossou /afrolm_active_learning_dataset AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages GitHub Repository of the Paper This repository contains the dataset for our paper AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages which will appear at the third Simple and Efficient Natural Language Processing, at EMNLP 2022. Our self-active learning framework Languages Covered AfroLM has been… See the full description on the dataset page: https://huggingface.co/datasets/bonadossou/afrolm_active_learning_dataset.imagefill-mask1M<n<10M5 likes902 downloads3y agoHugging Face13Ruinwalker /LabUtopia-Dataset🧪 LabUtopia-Dataset: Scientific Laboratory 3D Asset Library (OpenUSD) LabUtopia-Dataset is a large-scale 3D asset library designed for simulating scientific laboratory environments. It provides realistic lab scenes, scientific instruments, and environmental props, all stored in OpenUSD (.usd / .usdz) format for high interoperability and composability. 🧩 File Format: OpenUSD Each asset is stored as a .usd or .usdz file. You can load them directly in: NVIDIA Omniverse (Create, Isaac Sim)… See the full description on the dataset page: https://huggingface.co/datasets/Ruinwalker/LabUtopia-Dataset.imageroboticsn<1K0 likes859 downloads1y agoHugging Face14anonymousllbench /llbench-dataset LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models via Human Preferences Anonymous release prepared for NeurIPS 2026 review. Please do not redistribute. LL-Bench is a large-scale, human-preference benchmark for evaluating low-level vision restoration in the era of large generative models (LGMs). It compares 10 LGMs with 16 specilist and 5 all-in-one models across 16 low-level vision tasks, paired with dense human annotations:pairwise… See the full description on the dataset page: https://huggingface.co/datasets/anonymousllbench/llbench-dataset.imageimage-to-image100K<n<1M0 likes853 downloads4mo agoHugging Face15bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes734 downloads2y agoHugging Face16dartbrains /localizer Dartbrains Localizer Dataset A subset of the Brainomics/Localizer functional MRI dataset, prepared for the Dartbrains neuroimaging course at Dartmouth College. Quick Start Load beta maps (recommended for most exercises) from datasets import load_dataset ds = load_dataset("dartbrains/localizer", "betas") img = ds[0]["nifti"] # nibabel.Nifti1Image subject = ds[0]["subject"] # "S01" condition = ds[0]["condition"] # "audio_computation"… See the full description on the dataset page: https://huggingface.co/datasets/dartbrains/localizer.imageimage-classificationn<1K1 likes699 downloads3mo agoHugging Face17orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes608 downloads2y agoHugging Face18LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes605 downloads3mo agoHugging Face19Lo6yu /egocentric_dataset Egocentric RGB-D + EMG/IMU Daily Activity Dataset This dataset contains first-person daily activity recordings with synchronized RGB-D video, wrist EMG/IMU signals, hand keypoints, object masks, hand-object contact annotations, per-finger force annotations, and semantic action segments. Multimodal showcase video: RGB-D, hand joints, 3D hand projection, and object masks Overview The dataset is designed for egocentric embodied AI and robot learning in everyday… See the full description on the dataset page: https://huggingface.co/datasets/Lo6yu/egocentric_dataset.imageroboticsn<1K4 likes600 downloads3mo agoHugging Face20AvoCahDoe /llava-15-rlmpq-vlm-eval-results RL-MPQ VLM Evaluation Artifacts Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation. Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results Collections (by base VLM) RL-MPQ VLM — LLaVA-1.5-13B — HF collection RL-MPQ VLM — LLaVA-1.5-7B — HF collection RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection RL-MPQ VLM — Qwen2-VL-7B — HF collection Model repos RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.imagevisual-question-answeringn<1K0 likes529 downloads3mo agoHugging Face21Extend-AI /RealDoc-Bench-Layout RealDocBench-Layout A 1,500-page document-layout benchmark for evaluating layout-detection models on real-world documents. COCO-style annotations across 9 block classes. Contents images/ — 1,500 page images (PNG / JPG / occasional WebP-as-PNG; see Caveats). annotations/<pageId>.json — per-page COCO files, each with a single image record, an annotations list, a categories list, and a page_info block. manifest.csv — pageId → domain + source URLs. The canonical row… See the full description on the dataset page: https://huggingface.co/datasets/Extend-AI/RealDoc-Bench-Layout.imageobject-detection1K<n<10K5 likes468 downloads4mo agoHugging Face22Nininkkka /Minecraft-winter-shaders-Lite ❄️ Minecraft Winter Shaders (Lite) 10,000+ winter Minecraft screenshots with Photon Shaders. For generative models, style transfer, and high-quality vision tasks. Structure Single folder: winter/ Format: 960x540 JPEG License Dataset: CC BY NC 4.0.Shader: Photon Shaders by SixthSurge (commercial use of screenshots explicitly allowed).See PHOTON_SHADERS_LICENSE.txt. imageunconditional-image-generation1K<n<10K1 likes428 downloads7d agoHugging Face23letrinhan /vn-ebi-provincial Vietnam EBI provincial e-business index (2012-2026) VECOM E-Business Index (EBI) provincial scores, 14 report years (2012-2026. no official 2016). Subindexes HR&IT, B2C, B2B, and G2B when published. G2B dropped from the composite from 2021. Geographic panel is historical_63 through 2025 and current_34 (post-merger) in 2026. Weights and coverage vary by year - see year_meta. Chart digits verified against official PDFs/screenshots. published ebi_total kept even when a few VECOM… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-ebi-provincial.document1K<n<10K0 likes417 downloads3d agoHugging Face24jngb-labs /InvoiceBenchmark InvoiceBenchmark 200 synthetic invoices with cent-perfect ground truth, designed to measure the one thing language models are supposed to be able to do: read a number. The Pitch Invoice processing is the use case every enterprise AI pitch deck opens with. The numbers are either right or wrong, and the distance between right and wrong can be measured to the cent. This dataset exists because we ran the experiment and discovered that the gap between "this looks easy" and… See the full description on the dataset page: https://huggingface.co/datasets/jngb-labs/InvoiceBenchmark.documentquestion-answeringn<1K0 likes406 downloads5mo agoHugging Face25lampent /IRFL Dataset Card for IRFL Dataset Description Leaderboards Colab notebook code for IRFL evaluation Languages Dataset Structure Data Fields Dataset Creation Considerations for Using the Data Licensing Information Citation Information Dataset Description The IRFL dataset consists of idioms, similes, metaphors with matching figurative and literal images, and two novel tasks of multimodal figurative detection and retrieval.Using human annotation and an automatic pipeline… See the full description on the dataset page: https://huggingface.co/datasets/lampent/IRFL.image10K<n<100K5 likes399 downloads3y agoHugging Face26YinkaiW /LSV LSV: LabSuperVision Benchmark Dataset Description LSV is a multi-view video dataset of wet-lab biology experiments, captured from both first-person (XMglass smart glasses) and third-person (DJI action camera) perspectives. Each video records a researcher performing a laboratory protocol and is annotated with the corresponding protocol text, scene type, and—where applicable—deliberate procedural errors. The dataset is designed for research on: Protocol compliance… See the full description on the dataset page: https://huggingface.co/datasets/YinkaiW/LSV.imagevideo-classificationn<1K4 likes386 downloads6mo agoHugging Face27ty-li /Obstacle-Detection-Dataset-YOLO ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/ty-li/Obstacle-Detection-Dataset-YOLO.imageobject-detection10K<n<100K0 likes384 downloads15d agoHugging Face28lesc-unifi /beyond-the-brush Beyond the Brush: Fully-automated Crafting of Realistic Inpainted Images The generation of partially manipulated images is rapidly becoming a significant threat to the public's trust in online content. The proliferation of diffusion model-based tools that enable easy inpainting operations has significantly lowered the barrier to accessing these techniques. In this context, the multimedia forensics community finds itself at a disadvantage compared to attackers, as developing new… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/beyond-the-brush.imagemask-generation10K<n<100K0 likes382 downloads2y agoHugging Face29JdeRobot /Follow-Line-Combine-Datasetimage100K<n<1M0 likes366 downloads1y agoHugging Face30bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes353 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.