datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.CADS-dataset
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography
Overview
CADS is a robust, fully automated framework for segmenting 167 anatomical structures in Computed Tomography (CT), spanning from head to knee regions across diverse anatomical systems.
The framework consists of two main components:
CADS-dataset:
22,022 CT volumes with complete annotations for 167 anatomical structures.
Most extensive whole-body CT dataset… See the full description on the dataset page: https://huggingface.co/datasets/arekborucki/CADS-dataset.eu-ai-act-article-50-scoreboard
Article 50 historical public-evidence snapshot
This work was produced through an AI-assisted workflow directed by the author. Historical work used Anthropic assistance; the retrospective correction uses OpenAI GPT-6, with separate bounded Gemini advice. All three providers have products in the scored set.
Purpose: provide the corrected paper's version 1.1 bundle under v1_1. Start with its README and correction note. The paper and deposit and GitHub repository identify the same… See the full description on the dataset page: https://huggingface.co/datasets/NMAIResearch/eu-ai-act-article-50-scoreboard.MTBLS289
MTBLS289
A dataset of ~110 paired Whole Slide Images (WSI) and Mass Spectrometry Images (MSI).
Publication: Gerbig, S., Golf, O., Balog, J. et al. Analysis of colorectal adenocarcinoma tissue by desorption electrospray ionization mass spectrometric imaging. Anal Bioanal Chem 403, 2315–2325 (2012).
AraDiCE
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic.
As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.meshfleet_arenaassistive-ocr-data-acquisition
Assistive OCR Benchmark Data
Multilingual OCR benchmark for visually impaired assistance — Indian medicine labels, packaged goods, and signage in Bengali, Hindi, and English.
Dataset Summary
Property
Value
Total images
7,004 rows in manifest
Image sources
images/hf_medicines/, images/openfoodfacts/, images/synthetic/
Domains
medicine_packaging (6,588), packaged_goods (386), signage (30)
Languages
bn+en (5,337), hi+en (868), en (799)
Splits
dev… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/assistive-ocr-data-acquisition.Gutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.geolayers
Geolayers-Data
-->
This dataset card contains usage instructions and metadata for all data-products released with our paper:Using Multiple Input Modalities can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery. We release 3 modified versions of 3 benchmark datasets spanning land-cover segmentation, tree-cover regression, and multi-label land-cover classification tasks. These datasets are augmented with auxiliary, geographic inputs. A full list of… See the full description on the dataset page: https://huggingface.co/datasets/arjunrao2000/geolayers.skin-cancer-flagged-dataset
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/AreejAlotaibi12/skin-cancer-flagged-dataset.book-recommender-artifactsArtVision
README — ArtVision: Dataset per la valutazione delle competenze visivo-interpretative in dominio storico-artistico
Descrizione generale
Il dataset ArtVision è una raccolta di 250 task, organizzati in otto categorie, in cui immagini di repertori storico artisti realizzati tra il 1750 e il 1985, sono utilizzate come base per la costruzione di richieste a modelli multimodali. Il dataset permette di sviluppare un veloce test di valutazione di un modello multimodale… See the full description on the dataset page: https://huggingface.co/datasets/paolodegasperis/ArtVision.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.welsh-speech-3d-meshes
Welsh Speech Dataset - 3D Facial Meshes
3D facial reconstructions from the Welsh Speech Dataset.
Contents
3D meshes (.obj files) - One per frame
Texture maps (.png files) - Fused left-right stereo images from 3DMD
Captured using 3DMD 6-camera system
~330 zip files (one per speaker-phrase sequence)
File Structure
Files are organized as zip archives in the meshes/ directory, one zip per speaker-phrase sequence:
meshes/
├── speaker_01_phrase_01.zip
├──… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-3d-meshes.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.the-star-news-articlesMMQSD_ClipSyntel
Dataset Card for MMCQS Dataset
This is the MMCQS Dataset that have been used in the paper "CLIPSyntel: CLIP and LLM Synergy for Multimodal Question Summarization in Healthcare"
Github: https://github.com/AkashGhosh/CLIPSyntel-AAAI2024
Paper: https://arxiv.org/pdf/2312.11541
Uses
Download and unzip the Multimodal_images_finalnew.zip file, that can be found the in the 'Files and Version' section, to access the images that have been used in the dataset. The image… See the full description on the dataset page: https://huggingface.co/datasets/ArkaAcharya/MMQSD_ClipSyntel.quantarena-artifacts
QuantArena Artifact Bundle
Reproducibility artifacts for the paper QuantArena: Beat the Market or Be the
Market? A Live-Market Evaluation of Investment Paradigms (NeurIPS 2026
Evaluations & Datasets Track submission).
Summary
QuantArena is a controlled live-market evaluation protocol that holds the LLM
backend, market data stream, analyst workflow, capital, and execution harness
fixed across runs and varies only the investment doctrine (the policy
module). This bundle… See the full description on the dataset page: https://huggingface.co/datasets/NIPS26Repo/quantarena-artifacts.testMedical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.recipesccs_synthetic_translated_arabicThe columns inside the dataset as follows:
index
url
caption_en
caption_ar
The dataset size is 12556500 rows × 4 columns
ccs_synthetic_translated_arabic_processedsample_font_aesthetics_dsArabicConceptualCaptions3M
Arabic Translated Conceptual Captions Dataset
Overview
This dataset consists of conceptual captions translated into Arabic using the Google Translate API. It serves as a resource for researchers and developers interested in exploring the vision-language tasks and biases introduced during the translation process.
Dataset Information
Source Dataset: Conceptual Captions
Translation Tool: Google Translate API
Translation Language: English to Arabic… See the full description on the dataset page: https://huggingface.co/datasets/LinaAlhuri/ArabicConceptualCaptions3M.kaitz-retrieval-packbangla-news-articles-sampleM3Retrieve_IT2I
