datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Embodied-Captioning
Embodied Image Captioning – Manually Annotated Test Set
Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning
📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.us-layoffs-per-capita-by-state-warn-act
US layoffs per capita by state: WARN notices and affected workers per 100,000 residents, 2020-2026
Rebuilt 2026-09-24. 299 state-years across 48 states; 34,624 notices in the table.
Latest complete year 2025: District of Columbia leads at 5.302 notices per 100k residents
(36 notices, 4,268 workers reported); the median state is 0.744.
"Which states are losing the most jobs per resident?" is a question no state agency answers,
because each of them publishes only its own notices… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-per-capita-by-state-warn-act.conceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file.
We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs.
human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.DataComp_large_pool_BLIP2_captions
Dataset Card for DataComp_large_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_large_pool_BLIP2_captions.PRES
Personal and Relational Event Sequence Datasets
This dataset collection supports research on modeling personal and relational event sequences, as referred in our LoG 2025 paper: Integrating Sequential and Relational Modeling for User Events: Datasets and Prediction Tasks.
It is organized into two main folders: processed/ and tasks/.
Folder Structure
processed/
amazon-clothing_all_events.csv
amazon-electronics_all_events.csv
brightkite_all_events.csv… See the full description on the dataset page: https://huggingface.co/datasets/capitalone/PRES.captcha-csv-checkpointspolusa_capstonebao-val-coco-rating-cap
Dataset Summary
This dataset contains Japanese captions for COCO images and their English translations.The format is CSV.
Dataset Structure
Data Fields
The data fields are the same among all lines.
filename(str): The name of the COCO image file
chatgpt text(str): The text generated by gpt-5-pro
gemini text(str): The text generated by gemini-3-pro-preview
grok text(str): The text generated by grok-4
claude text(str): The text generated by… See the full description on the dataset page: https://huggingface.co/datasets/baobab-trees/bao-val-coco-rating-cap.aav2_capsid_viability
AAV2 Capsid Viability Dataset
This dataset contains a preprocessed version of the Adeno-associated virus 2 (AAV2) capsid viability dataset from Bryant et al. 2021, including the full VP1, VP2, and VP3 sequences for each variant.
Description
This processed version of the dataset contains 289,805 variants with the following columns:
variable_region_sequence (str): The unique amino acid sequence of the variable region for each AAV2 variant
source_partition (str):… See the full description on the dataset page: https://huggingface.co/datasets/bviggiano/aav2_capsid_viability.vn-provinces-grdp-per-capita
Vietnam provinces GRDP per capita
Provincial and regional GRDP per capita (million VND per person). Coverage 2018-2024. Year 2024 is preliminary. Tables cover provinces, regions. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (441 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (42… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-grdp-per-capita.capterra-b2b-software-reviews
Capterra B2B Software Reviews
56,606 B2B software reviews from Capterra, covering 66 products across 11 software categories.
Most public review datasets are star rating + review text. This one carries five separate rating dimensions, pros and cons as distinct pre-split fields, reviewer firmographics, and, unusually, an incentive disclosure flag recording whether the reviewer was given a gift card, referred by the vendor, or wrote the review unprompted.
Why this is… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/capterra-b2b-software-reviews.vn-provinces-cereal-production-per-capita
Vietnam provinces cereal production per capita
Provincial and regional cereal production per capita (kg per person). Coverage 1995-2024. Year 2024 is preliminary. Includes historical Ha Tay through 2007 (dissolved into Ha Noi). Tables cover provinces, regions and national total. Geographic labels are English (UN/GSO style ASCII romanization). Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-cereal-production-per-capita.vn-provinces-enterprise-average-capital
Vietnam provinces enterprise average capital
Average annual production and business capital of operating enterprises with business results. Coverage 2010, 2015-2023. Billion VND. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (630 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-enterprise-average-capital.vn-provinces-housing-area-per-capita-by-type
Vietnam housing area per capita by house type
Vietnam housing area per capita by house type. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (378 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (36 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-housing-area-per-capita-by-type.vn-provinces-aquatic-capture-production
Vietnam provinces aquatic capture production
Capture fisheries production (thousand tons). Source values in tons converted to thousand tons. Coverage 1995-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-aquatic-capture-production.vn-provinces-income-per-capita-by-quintile
Vietnam income per capita by quintile
Vietnam income per capita by quintile. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (693 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (66 rows)
data/regions.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-income-per-capita-by-quintile.DataComp_medium_pool_BLIP2_captions
Dataset Card for DataComp_medium_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_medium_pool_BLIP2_captions.Full_Length_Motion_Capture_Dataset
Apple Arts Studios Full-Length Motion Capture Dataset
Dataset Overview
The Apple Arts Studios Full-Length Motion Capture Dataset is a professionally captured, full-body human-motion dataset containing 199 hours and 30 minutes of continuous motion capture data.
Unlike segmented motion datasets, this repository preserves the complete capture sequences without separating individual actions into short clips.
The recordings retain their continuous capture… See the full description on the dataset page: https://huggingface.co/datasets/Appleartsstudios/Full_Length_Motion_Capture_Dataset.vn-provinces-marine-fish-capture-production
Vietnam provinces marine fish capture production
Marine fish capture production (thousand tons). Coastal provinces. Coverage varies by year. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Hero (continued)
Comparison
Color key
Files
provinces (870 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-marine-fish-capture-production.length_captchas
Length Captchas Dataset
Converted dataset from twitter_ai label tool.
Nagi_no_Asukara_Videos_Captioned
Reorganized version of Wild-Heart/Disney-VideoGeneration-Dataset. This is needed for Mochi-1 fine-tuning.
dreamlip_long_captions
Dataset Card for DreamLIP-30M
Dataset Summary
DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.vn-provinces-enterprises-by-capital-size
Vietnam provinces enterprises by capital size
Operating enterprises with business results by registered-capital bins (tỷ đồng) as of 31 December. Coverage 2021-2023. Size bins sum to total. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (189… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-enterprises-by-capital-size.DATASET-CAPE-RhlA-seqlabel
CAPE Dataset: RhlA Enzyme Mutations(https://github.com/KRATSZ/CAPE-2023/tree/main)
Dataset Introduction and Usage 😋
RhlA (Uniprot ID: Q51559, PDB ID: 8IK2) is a key enzyme involved in synthesizing the hydrophobic component of rhamnolipids. The enzyme determines the length and unsaturation of the fatty acid chains, which ultimately influences the physicochemical properties and bioactivity of rhamnolipids.
Protein Format: AA sequence
Why Modify RhlA?
Modifying… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/DATASET-CAPE-RhlA-seqlabel.circuitkit-capitals-contrastive
CircuitKIT capitals — contrastive pairs
Twelve capital-city facts, each with an explicit counterfactual pair, for circuit
discovery with CircuitKIT.
column
meaning
question
clean prompt, e.g. The capital of France is
answer
clean answer, e.g. Paris
corrupted_question
counterfactual prompt of the same shape, e.g. The capital of Germany is
corrupted_answer
its answer, e.g. Berlin
Attribution-patching methods (EAP, EAP-IG, …) score a component by how much it… See the full description on the dataset page: https://huggingface.co/datasets/Lexsi/circuitkit-capitals-contrastive.GDP-Per-Capita_Gov-Expenditure_TradeAbout Dataset
Edit
This dataset provides a comprehensive view of key macroeconomic indicators across various entities (countries or regions) over time. It includes annual data for the following variables:
Entity: The name of the country or region for which the data is recorded.
Code: A standardized three-letter country or region code, facilitating easier identification and merging with other datasets.
Year: The calendar year for which the economic indicators are reported.
GDP per capita:… See the full description on the dataset page: https://huggingface.co/datasets/tripathyShaswata/GDP-Per-Capita_Gov-Expenditure_Trade.capstone_testCAP-Bench
CAP-Bench
A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception.
CAP-Bench evaluates browser agents on Cross-site workflows, complex Actions, and challenging visual Perception. The full benchmark contains 420 tasks across 108 real-world websites in 24 functional domains. Each task requires on average 7 complex execution operations and 4 perception challenges, substantially exceeding the difficulty of prior browser-agent benchmarks.
This… See the full description on the dataset page: https://huggingface.co/datasets/Warrior0302/CAP-Bench.
