CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01imolodetskikh /sr-artifact-prominence SR Artifact Prominence Annotated super-resolution artifact regions across four image subsets, with crowdsourced per-region prominence scores, artifact type labels, and natural-language descriptions. Prominence is the fraction of valid crowd workers who answered that the highlighted region contains a noticeable super-resolution artifact. Subsets Subset Source dataset Source images Masks Notes open_images Open Images 547 1,523 GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.image1K<n<10K0 likes2.3k downloads5mo agoHugging Face02aditya487 /cbi-archive-raw Central Bank of Ireland Archive: original source files 6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and archive gathered from the Central Bank of Ireland's public website, stored by content hash so that a search result can be turned back into the document a human would actually read. This is the raw tier. If you want the text, you almost certainly want aditya487/cbi-archive-corpus instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.document1K<n<10K0 likes1.3k downloads25d agoHugging Face03arekborucki /CADS-dataset CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems. The framework consists of two main components: CADS-dataset: 22,022 CT volumes with complete annotations for 167 anatomical structures. Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/arekborucki/CADS-dataset.tabularimage-segmentation10K<n<100K2 likes947 downloads9mo agoHugging Face04NMAIResearch /eu-ai-act-article-50-scoreboard Article 50 historical public-evidence snapshot This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set. Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.imagen<1K0 likes835 downloads3d agoHugging Face05aryan-f /MTBLS289 MTBLS289 A dataset of ~110 paired Whole Slide Images (WSI) and Mass Spectrometry Images (MSI). Publication: Gerbig, S., Golf, O., Balog, J. et al. Analysis of colorectal adenocarcinoma tissue by desorption electrospray ionization mass spectrometric imaging. Anal Bioanal Chem 403, 2315–2325 (2012). imageimage-to-imagen<1K0 likes299 downloads2mo agoHugging Face06QCRI /AraDiCE AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.imagetext-classificationn<1K3 likes219 downloads2y agoHugging Face07QTE-Technologies /industrial-technical-archive 🚀 Latest Updates (July, 2026) Version: v07.2026 (Verified) Status: Integrated with 1,000,000+ records. New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv. QTE Technologies: Industrial & Scientific Knowledge Base Wikidata Entity: Q138411149 IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq Official Neural Hub: qtetech.github.io This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.image10K<n<100K0 likes218 downloads2mo agoHugging Face08DamianBoborzi /meshfleet_arena3dn<1K0 likes188 downloads4mo agoHugging Face09Arko007 /assistive-ocr-data-acquisition Assistive OCR Benchmark Data Multilingual OCR benchmark for visually impaired assistance — Indian medicine labels, packaged goods, and signage in Bengali, Hindi, and English. Dataset Summary Property Value Total images 7,004 rows in manifest Image sources images/hf_medicines/, images/openfoodfacts/, images/synthetic/ Domains medicine_packaging (6,588), packaged_goods (386), signage (30) Languages bn+en (5,337), hi+en (868), en (799) Splits dev… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/assistive-ocr-data-acquisition.imageimage-to-text1K<n<10K0 likes136 downloads2mo agoHugging Face10FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes119 downloads1y agoHugging Face11arjunrao2000 /geolayers Geolayers-Data --> This dataset card contains usage instructions and metadata for all data-products released with our paper:Using Multiple Input Modalities can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery. We release 3 modified versions of 3 benchmark datasets spanning land-cover segmentation, tree-cover regression, and multi-label land-cover classification tasks. These datasets are augmented with auxiliary, geographic inputs. A full list of… See the full description on the dataset page: https://huggingface.co/datasets/arjunrao2000/geolayers.imageimage-classificationn<1K0 likes96 downloads1y agoHugging Face12AreejAlotaibi12 /skin-cancer-flagged-dataset Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/AreejAlotaibi12/skin-cancer-flagged-dataset.imagen<1K0 likes76 downloads2y agoHugging Face13nice-bill /book-recommender-artifactsimage10K<n<100K0 likes73 downloads10mo agoHugging Face14paolodegasperis /ArtVision README — ArtVision: Dataset per la valutazione delle competenze visivo-interpretative in dominio storico-artistico Descrizione generale Il dataset ArtVision è una raccolta di 250 task, organizzati in otto categorie, in cui immagini di repertori storico artisti realizzati tra il 1750 e il 1985, sono utilizzate come base per la costruzione di richieste a modelli multimodali. Il dataset permette di sviluppare un veloce test di valutazione di un modello multimodale… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/ArtVision.imagen<1K0 likes61 downloads7mo agoHugging Face15crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes38 downloads6mo agoHugging Face16arvinsingh /welsh-speech-3d-meshes Welsh Speech Dataset - 3D Facial Meshes 3D facial reconstructions from the Welsh Speech Dataset. Contents 3D meshes (.obj files) - One per frame Texture maps (.png files) - Fused left-right stereo images from 3DMD Captured using 3DMD 6-camera system ~330 zip files (one per speaker-phrase sequence) File Structure Files are organized as zip archives in the meshes/ directory, one zip per speaker-phrase sequence: meshes/ ├── speaker_01_phrase_01.zip ├──… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-3d-meshes.3dimage-to-3dn<1K0 likes32 downloads8mo agoHugging Face17crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes30 downloads1y agoHugging Face18azrai99 /the-star-news-articlesimage10K<n<100K1 likes22 downloads2y agoHugging Face19ArkaAcharya /MMQSD_ClipSyntel Dataset Card for MMCQS Dataset This is the MMCQS Dataset that have been used in the paper "CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare" Github: https://github.com/AkashGhosh/CLIPSyntel-AAAI2024 Paper: https://arxiv.org/pdf/2312.11541 Uses Download and unzip the Multimodal_images_finalnew.zip file, that can be found the in the 'Files and Version' section, to access the images that have been used in the dataset. The image… See the full description on the dataset page: https://huggingface.co/datasets/ArkaAcharya/MMQSD_ClipSyntel.imagesummarization1K<n<10K2 likes19 downloads2y agoHugging Face20NIPS26Repo /quantarena-artifacts QuantArena Artifact Bundle Reproducibility artifacts for the paper QuantArena: Beat the Market or Be the Market? A Live-Market Evaluation of Investment Paradigms (NeurIPS 2026 Evaluations & Datasets Track submission). Summary QuantArena is a controlled live-market evaluation protocol that holds the LLM backend, market data stream, analyst workflow, capital, and execution harness fixed across runs and varies only the investment doctrine (the policy module). This bundle… See the full description on the dataset page: https://huggingface.co/datasets/NIPS26Repo/quantarena-artifacts.imagetabular-classification10K<n<100K1 likes19 downloads5mo agoHugging Face21Arabic-Image-Captioning-latest /testimage1M<n<10M2 likes17 downloads3y agoHugging Face22david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes17 downloads4mo agoHugging Face23arya123321 /recipesimage10K<n<100K4 likes14 downloads3y agoHugging Face24Arabic-Clip /ccs_synthetic_translated_arabicThe columns inside the dataset as follows: index url caption_en caption_ar The dataset size is 12556500 rows × 4 columns image10M<n<100M0 likes13 downloads2y agoHugging Face25Arabic-Clip /ccs_synthetic_translated_arabic_processedimage10M<n<100M1 likes13 downloads2y agoHugging Face26arpitdvd /sample_font_aesthetics_dsimagen<1K0 likes10 downloads3y agoHugging Face27LinaAlhuri /ArabicConceptualCaptions3M Arabic Translated Conceptual Captions Dataset Overview This dataset consists of conceptual captions translated into Arabic using the Google Translate API. It serves as a resource for researchers and developers interested in exploring the vision-language tasks and biases introduced during the translation process. Dataset Information Source Dataset: Conceptual Captions Translation Tool: Google Translate API Translation Language: English to Arabic… See the full description on the dataset page: https://huggingface.co/datasets/LinaAlhuri/ArabicConceptualCaptions3M.imageimage-to-text1M<n<10M3 likes9 downloads3y agoHugging Face28arielgd30 /kaitz-retrieval-packimage1K<n<10K0 likes9 downloads8mo agoHugging Face29shmazumder-cse /bangla-news-articles-sampleimagen<1K0 likes8 downloads2y agoHugging Face30ArkaAcharya /M3Retrieve_IT2Iimage1K<n<10K0 likes5 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.