datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.TreeOfLife-10M-Captions
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M-Captions.pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
CAPTURe
CAPTURe Dataset
This is the dataset for CAPTURe, a new benchmark and task to evaluate spatial reasoning in vision-language models, as described in the paper:
CAPTURE: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
by Atin Pothiraj, Elias Stengel-Eskin, Jaemin Cho, Mohit Bansal
Code is available here.
Overview
Recognizing and reasoning about occluded (partially or fully hidden) objects is vital to understanding visual scenes, as… See the full description on the dataset page: https://huggingface.co/datasets/atinp/CAPTURe.Products-10k-BLIP-captions
Dataset Description
The Products-10k BLIP CAPTIONS dataset consists of 10000 images of various products along with their automatically generated captions. The captions are generated using the BLIP (Bootstrapping Language-Image Pre-training) model. This dataset aims to aid in tasks related to image captioning, visual recognition, and product classification.
Dataset Summary
Dataset Name: Products-10k
Generated Captions Model: Salesforce/blip-image-captioning-large… See the full description on the dataset page: https://huggingface.co/datasets/VikramSingh178/Products-10k-BLIP-captions.captchaimages
