any
Datasets
All datasets matching “any”AnyAudio-Judge-Corpus
AnyAudio-Judge Corpus
An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains:
An audio clip (referenced relatively under audios/).
A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale).
A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.anypoint-2map-anything
MapAnything Training Metadata Dataset
Dataset Description
This dataset contains pre-computed metadata and covisibility matrices for supporting the MapAnything codebase. This metadata enables easy reproducible training for feed-forward 3D reconstruction tasks.
Please see our Data Processing README for more details.
Citation
If you use this dataset in your research, please cite our paper:
@inproceedings{keetha2026mapanything,
title={{MapAnything}: Universal… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything.AnyWord-3MDataset from AnyText: Multilingual Visual Text Generation And Editing.
Dataset description from Anytext Team:
Currently, there is a relative scarcity of public datasets for text generation tasks, especially those involving non-Latin script languages. To address this, we introduce a large-scale multilingual dataset called AnyWord-3M. The images in this dataset are sourced from Noah-Wukong, LAION-400M, and OCR recognition datasets such as ArT, COCO-Text, RCTW, LSVT, MLT, MTWI, ReCTS, etc. These… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/AnyWord-3M.Grasp-Any-Region-Dataset
Grasp Any Region Dataset
This repository contains the training dataset for the paper: Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs.
Code: https://github.com/Haochen-Wang409/Grasp-Any-Region
About the Dataset
The Grasp Any Region (GAR) dataset is designed to empower Multimodal Large Language Models (MLLMs) with comprehensive region-level visual understanding. While MLLMs excel at holistic understanding, they often struggle with… See the full description on the dataset page: https://huggingface.co/datasets/HaochenWang/Grasp-Any-Region-Dataset.align-anything
Overview: Align-Anything Dataset
A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback.
🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo
Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.
