datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.anime-art-curated
anime-art-curated
A curated WebDataset of anime / digital illustration images with Danbooru-style
text tags, intended as a clean training corpus for text-to-image and
illustration-style model fine-tuning.
Provenance
This dataset is a filtered re-publication of an earlier dataset
(advokat/artist390k, now deleted) which was itself algorithmically curated
from broader booru sources by selecting popular images. The original source
inadvertently contained material that the… See the full description on the dataset page: https://huggingface.co/datasets/advokat/anime-art-curated.sdxl_images_easy_prompts-artists-seed1Argimi-Ardian-Finance-10k-text-image
The ArGiMI Ardian datasets : text and images
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural language… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text-image.art-museums-pd-440k
Art Museums PD 440K
Summary
This is a dataset to train text-to-image or any text and image multimodal models with minimized copyright/licensing concerns.
All images and texts in this dataset are orignally shared under CC0 or public domain, and no pretrained models or any AI models are used to build this dataset except for our ElanMT model to translate English captions to Japanese.
ElanMT model is trained solely on licensed corpus.
Data sources
Images and… See the full description on the dataset page: https://huggingface.co/datasets/Mitsua/art-museums-pd-440k.infinigen-articulated
Infinigen-Articulated Assets
Formerly Infinigen-Sim
sdxl_images_sb_prompts-multi_artist-seed1physical-ai-bench-artifactsdtasettarsphere-encoder-fid-artifacts
Sphere Encoder FID Evaluation Artifacts
This repository contains the evaluation artifacts for the paper Image Generation with a Sphere Encoder.
Project Page | GitHub Repository
These artifacts include data statistic files (fid_stats) and reference images (fid_refs) used to calculate Fréchet Inception Distance (FID) for generative models across several datasets, including CIFAR-10, ImageNet, Animal Faces, and Oxford Flowers.
Workspace Setup
Download the evaluation… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/sphere-encoder-fid-artifacts.dtasettar23Garments2Look-Test-Set-Results
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories
Project Page | Paper | Code
Garments2Look is a large-scale multimodal dataset for outfit-level Virtual Try-On (VTON), comprising 80,000 many-garments-to-one-look pairs across 40 major categories and over 300 fine-grained subcategories. Each pair includes an outfit with 3-12 reference garment images (averaging 4.48), a model image wearing the outfit, and detailed item… See the full description on the dataset page: https://huggingface.co/datasets/ArtmeScienceLab/Garments2Look-Test-Set-Results.twoframe-eval-artifacts-20260505
TwoFrame Eval Artifacts 2026-05-05
Generated image artifacts for TwoFrame image-editing evaluation. The large image payload is stored as tar archives under archives/ to avoid uploading tens of thousands of loose PNG files.
Layout
archives/single_ref.tar: all complete single-reference outputs.
archives/multiref_part*.tar: complete K=2/K=3 multi-reference runs, sharded by run name.
archives/metadata.tar: README, index, manifests, and metrics as an archive.… See the full description on the dataset page: https://huggingface.co/datasets/wyhhey/twoframe-eval-artifacts-20260505.ArTarticubot
ArticuBot Simulated Dataset for Trajectories of Articulation (saved in WebDataset Format and splitter into train / val)
Data Format
Each sample in the WebDataset contains:
metadata.json: Complete metadata including object_id, category, and trajectory info
trajectory.pkl: Trajectory data with all timesteps
info.json: Sample structure information
Sample Structure
Each trajectory contains timesteps with the following data:
state: Robot state information
action:… See the full description on the dataset page: https://huggingface.co/datasets/LocalWorldModels/articubot.ArtifactWorld-Benchmarkmjnj_tarjuejin_article_introsdxl_images_easy_prompts-multi_artist-seed0openpi-mcts-artifactssdxl_images_sb_prompts-multi_artist-seed0sdxl_images_easy_prompts-multi_artist-seed1sdxl_images_easy_prompts-artists-seed0turboquant-art-5kturboquant-art-1kturboquant-art-50kartgrid-human
