datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AnyWord-3MDataset from AnyText: Multilingual Visual Text Generation And Editing.
Dataset description from Anytext Team:
Currently, there is a relative scarcity of public datasets for text generation tasks, especially those involving non-Latin script languages. To address this, we introduce a large-scale multilingual dataset called AnyWord-3M. The images in this dataset are sourced from Noah-Wukong, LAION-400M, and OCR recognition datasets such as ArT, COCO-Text, RCTW, LSVT, MLT, MTWI, ReCTS, etc. These… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/AnyWord-3M.Grasp-Any-Region-Dataset
Grasp Any Region Dataset
This repository contains the training dataset for the paper: Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs.
Code: https://github.com/Haochen-Wang409/Grasp-Any-Region
About the Dataset
The Grasp Any Region (GAR) dataset is designed to empower Multimodal Large Language Models (MLLMs) with comprehensive region-level visual understanding. While MLLMs excel at holistic understanding, they often struggle with… See the full description on the dataset page: https://huggingface.co/datasets/HaochenWang/Grasp-Any-Region-Dataset.align-anything
Overview: Align-Anything Dataset
A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback.
🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo
Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.AnyAudio-Judge-Bench
AnyAudio-Judge Bench
Bilingual (English / Chinese) multi-domain benchmark for instruction-audio alignment evaluation, released alongside the paper "AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following".
7,920 curated samples per language across 7 subsets
Strict 1 : 1 positive : negative ratio per subset
Hard negatives via instruction swapping and attribute perturbation
Each row carries a list of decomposed binary rubric items (yes/no… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Bench.AnyEdit
Celebrate! AnyEdit resolved the data alignment with the re-uploading process (but the view filter is not working:(, though it has 25 edit types). You can view the validation split for a quick look. You can also refer to anyedit-split dataset to view and download specific data for each editing type.
Dataset Card for AnyEdit-Dataset
Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often… See the full description on the dataset page: https://huggingface.co/datasets/Bin1117/AnyEdit.anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.any4hdmi-g1-100styledescribe-anything-dataset
Describe Anything: Detailed Localized Image and Video Captioning
NVIDIA, UC Berkeley, UCSF
Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui
[Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation]
Dataset Card for Describe Anything Datasets
Datasets used in the training of describe anything models (DAM).
The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.DA3-BENCH
DA3-BENCH: Depth Anything 3 Evaluation Benchmark
This repository contains processed benchmark datasets for evaluating Depth Anything 3 depth estimation and visual geometry models. The datasets are provided in a convenient, ready-to-use format for research and evaluation purposes.
About Depth Anything 3
Depth Anything 3 (DA3) is a state-of-the-art model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known… See the full description on the dataset page: https://huggingface.co/datasets/depth-anything/DA3-BENCH.anyedit-splitAnyInsertion
AnyInsertion
Wensong Song
·
Hong Jiang
·
Zongxing Yang
·
Ruijie Quan
·
Yi Yang
Zhejiang University | Harvard University | Nanyang Technological University
News
[2025.5.9] Release new AnyInsertion v1 text- and mask-prompt dataset on HuggingFace.
[2025.5.7] Release AnyInsertion v1 text prompt dataset on HuggingFace.
[2025.4.24] Release AnyInsertion v1 mask prompt dataset on HuggingFace.
Summary
This is the dataset proposed in… See the full description on the dataset page: https://huggingface.co/datasets/WensongSong/AnyInsertion.AnyPatternThe dataset proposed in our paper "AnyPattern: Towards In-context Image Copy Detection".
Please go to Github for the code about how to use this dataset.
Here, we show how to download this dataset.
anypattern_v31
for letter in {a..z}; do
wget https://huggingface.co/datasets/WenhaoWang/AnyPattern/resolve/main/train/anypattern_v31_part_a$letter
done
wget https://huggingface.co/datasets/WenhaoWang/AnyPattern/resolve/main/train/anypattern_v31_part_ba
cat anypattern_v31_part_a{a..z}… See the full description on the dataset page: https://huggingface.co/datasets/WenhaoWang/AnyPattern.ipapack_plus_train_3AnyInsertion_V1
AnyInsertion
Wensong Song
·
Hong Jiang
·
Zongxing Yang
·
Ruijie Quan
·
Yi Yang
Zhejiang University | Harvard University | Nanyang Technological University
News
[2025.5.9] Release new AnyInsertion v1 text- and mask-prompt dataset on HuggingFace.
[2025.5.7] Release AnyInsertion v1 text prompt dataset on HuggingFace.
[2025.4.24] Release AnyInsertion v1 mask prompt dataset on HuggingFace.
Summary
This is the dataset proposed in… See the full description on the dataset page: https://huggingface.co/datasets/WensongSong/AnyInsertion_V1.dualturn-otospeech-turn-taking
OtoSpeech Turn-Taking
Official DualTurn release of the otospeech corpus, with per-frame turn-taking labels and
Mimi speech codec features. Each row is one full conversation. Frame rate 12.5 Hz (80 ms per frame).
Paper: DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
Training code: github.com/anyreachai/dualturn
Model checkpoint: anyreach-ai/dualturn-qwen2.5-mimi-0.5B
Splits
Split
Sessions
train
896
val
111
test
113… See the full description on the dataset page: https://huggingface.co/datasets/anyreach-ai/dualturn-otospeech-turn-taking.waxal-pseudo
WAXAL Pseudo-Labels (3-model agreement)
PRIVATE working artifact for the Google WAXAL ASR Challenge — not for redistribution.
High-confidence pseudo-labels for the WAXAL unlabeled split, produced by a 3-model agreement cascade:
a clip is kept only when the fine-tuned champion (w2v-BERT-2.0 CTC) and XLS-R-300m agree
(CER ≤ 0.12), and the fine-tuned Omnilingual-ASR-300M independently confirms the champion transcript
(CER ≤ 0.22). omni is architecturally diverse (different… See the full description on the dataset page: https://huggingface.co/datasets/anyantudre/waxal-pseudo.Align-Anything-Cosiucla_phonetic_corpus
Dataset Card for "ucla_phonetic_corpus"
More Information needed
AnyMo-Bench
AnyMo Bench
AnyMo Bench is a challenging fine-grained in-the-wild HAR benchmark built from real wearable IMU streams in the Nymeria dataset. It provides unseen-subject and cross-device evaluation settings for wearable motion recognition.
For general project information, see the AnyMo project page. For more technical details, see the AnyMo paper. The code is available at Breezelled/AnyMo.
The benchmark contains 154,695 eligible activity windows from 196 subjects, covering 211.6… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/AnyMo-Bench.synthetic-driver-monitoring-detection
Synthetic DMS – Driver Monitoring System Dataset by AnywayLabs.ai
Need a custom synthetic dataset for your own road safety detection use case?
This dataset is an open-source sample of our synthetic data generation work at AnywayLabs.
If you're working on:
industrial defect detection
visual inspection
supervised anomaly detection
hard-to-collect defect classes
synthetic data for computer vision training
You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-driver-monitoring-detection.anyeditflickr30kany4hdmi-g1-lafanAlign-Anything-L0pixmo-ask-model-anything
PixMo-AskModelAnything
PixMo-AskModelAnything is an instruction-tuning dataset for vision-language models. It contains human-authored
question-answer pairs about diverse images with long-form answers.
PixMo-AskModelAnything is a part of the PixMo dataset collection and was used to train the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with Videos
Loading
data = datasets.load_dataset("allenai/pixmo-ask-model-anything", split="train")
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-ask-model-anything.Align-Anything-Coccurdocument-qna-chroma-anyscale-logsdualturn-switchboard-turn-taking
Switchboard Turn-Taking
Official DualTurn release of the switchboard corpus, with per-frame turn-taking labels and
Mimi speech codec features. Each row is one full conversation. Frame rate 12.5 Hz (80 ms per frame).
Paper: DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
Training code: github.com/anyreachai/dualturn
Model checkpoint: anyreach-ai/dualturn-qwen2.5-mimi-0.5B
Splits
Split
Sessions
train
1986
val
295
test
138… See the full description on the dataset page: https://huggingface.co/datasets/anyreach-ai/dualturn-switchboard-turn-taking.MilitaryAircraftRecognitionsynthetic-mvtec-ad-defect-detection
Synthetic MVTec AD – Defect Detection Dataset by AnywayLabs.ai
Need a custom synthetic dataset for your own defect detection use case?
This dataset is an open-source sample of our synthetic data generation work at AnywayLabs.
If you're working on:
industrial defect detection
visual inspection
supervised anomaly detection
hard-to-collect defect classes
synthetic data for computer vision training
You can request a custom synthetic dataset here, or email:… See the full description on the dataset page: https://huggingface.co/datasets/anywaylabs/synthetic-mvtec-ad-defect-detection.
