CoolFace
20 results

any

cucl2 /AnyAudio-Judge-Corpus AnyAudio-Judge Corpus An SFT training corpus that powers the AnyAudio-Judge evaluator. Each sample contains: An audio clip (referenced relatively under audios/). A multi-turn chat (messages) where the user enumerates a list of decomposed binary rubric items and the assistant answers them in JSON, with per-item evidence (Chain-of-Thought rationale). A coarse label ("yes" if the caption originally matched the audio, "no" otherwise) and a tag describing how the caption was… See the full description on the dataset page: https://huggingface.co/datasets/cucl2/AnyAudio-Judge-Corpus.audioaudio-text-to-text10K<n<100K0 likes11k downloads2mo agoHugging FaceKoios /anypoint-20 likes10k downloads7mo agoHugging Facefacebook /map-anything MapAnything Training Metadata Dataset Dataset Description This dataset contains pre-computed metadata and covisibility matrices for supporting the MapAnything codebase. This metadata enables easy reproducible training for feed-forward 3D reconstruction tasks. Please see our Data Processing README for more details. Citation If you use this dataset in your research, please cite our paper: @inproceedings{keetha2026mapanything, title={{MapAnything}: Universal… See the full description on the dataset page: https://huggingface.co/datasets/facebook/map-anything.image-to-3d100B<n<1T9 likes7.8k downloads9mo agoHugging Facestzhao /AnyWord-3MDataset from AnyText: Multilingual Visual Text Generation And Editing. Dataset description from Anytext Team: Currently, there is a relative scarcity of public datasets for text generation tasks, especially those involving non-Latin script languages. To address this, we introduce a large-scale multilingual dataset called AnyWord-3M. The images in this dataset are sourced from Noah-Wukong, LAION-400M, and OCR recognition datasets such as ArT, COCO-Text, RCTW, LSVT, MLT, MTWI, ReCTS, etc. These… See the full description on the dataset page: https://huggingface.co/datasets/stzhao/AnyWord-3M.imagetext-to-image1M<n<10M17 likes7k downloads2y agoHugging FaceHaochenWang /Grasp-Any-Region-Dataset Grasp Any Region Dataset This repository contains the training dataset for the paper: Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs. Code: https://github.com/Haochen-Wang409/Grasp-Any-Region About the Dataset The Grasp Any Region (GAR) dataset is designed to empower Multimodal Large Language Models (MLLMs) with comprehensive region-level visual understanding. While MLLMs excel at holistic understanding, they often struggle with… See the full description on the dataset page: https://huggingface.co/datasets/HaochenWang/Grasp-Any-Region-Dataset.textimage-text-to-text1M<n<10M3 likes6.9k downloads11mo agoHugging FacePKU-Alignment /align-anything Overview: Align-Anything Dataset A Comprehensive All-Modality Alignment Dataset with Fine-grained Preference Annotations and Language Feedback. 🏠 Homepage | 🤗 Align-Anything Dataset | 🤗 T2T_Instruction-tuning Dataset | 🤗 TI2T_Instruction-tuning Dataset | 👍 Our Official Code Repo Our world is inherently multimodal. Humans perceive the world through multiple senses, and Language Models should operate similarly. However, the development of Current Multi-Modality Foundation Models… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/align-anything.audioany-to-any10K<n<100K48 likes6.3k downloads1y agoHugging Face