datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.OmniPhysics-Caption_benchmark
OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning
🌐 Project •
📈 Benchmark Overview •
🧪 What OPC Measures •
📊 Daily-Physics 50K Subset •
📚 Citation
Introduction
Building omni-modal models with physical intelligence requires benchmarks that
test whether generated captions preserve information from both visual and
audio streams. Existing detailed-caption benchmarks provide strong visual… See the full description on the dataset page: https://huggingface.co/datasets/Fysics-AI/OmniPhysics-Caption_benchmark.TreeOfLife-10M-Captions
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M-Captions.LiDAR-LLM-Nu-Caption
Dataset Details
Dataset type:
This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset.
Dataset keys:
"answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data.
If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI.
Products-10k-BLIP-captions
Dataset Description
The Products-10k BLIP CAPTIONS dataset consists of 10000 images of various products along with their automatically generated captions. The captions are generated using the BLIP (Bootstrapping Language-Image Pre-training) model. This dataset aims to aid in tasks related to image captioning, visual recognition, and product classification.
Dataset Summary
Dataset Name: Products-10k
Generated Captions Model: Salesforce/blip-image-captioning-large… See the full description on the dataset page: https://huggingface.co/datasets/VikramSingh178/Products-10k-BLIP-captions.LISA_Plus_Caption
LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
🤗Data | 📄Paper |
🚀Code | 💻Model |
🔥Citation
Dataset Details
Dataset type:
The LISA++ Caption dataset is a QA dataset designed to train MLLM models for segmentation in captioning. It is based on the COCO2017 dataset.
Where to send questions or comments about the dataset:
https://github.com/dvlab-research/LISA
Paper:https://arxiv.org/abs/2312.17240
This model could be used for… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LISA_Plus_Caption.flickr30k-captions_marathi
Flickr30K-Captions Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Flickr30K-Captions Marathi dataset is a meticulously curated collection of 158881 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/flickr30k-captions_marathi.coco-captions_marathi
Coco-Captions Marathi Dataset: High-Quality Marathi NLP Corpus
📌 Overview
The Coco-Captions Marathi dataset is a meticulously curated collection of 414010 rows of Marathi text, ensuring linguistic accuracy and natural flow. Every sentence has been verified by native Marathi speakers to maintain contextual integrity and correctness.
This dataset is designed for semantic search, text classification, and various NLP tasks, making it a valuable resource for machine learning… See the full description on the dataset page: https://huggingface.co/datasets/Singhchandann/coco-captions_marathi.refined-anime-instruct-en-641k
Dataset Card for refined-anime-instruct-en-641k
Dataset Summary
This is 641,497 instructions for an expert model that knows about the following things:
Anime
Manga
Live Action Shows
Children's Films
Western Comics
Agatha Christie Novels and Adaptations (not sure why this is over-represented)
Video Games
It is derived from Refined-Anime-Text by filtering out all ZH entries. According to their README.md, these outputs are completions derived from GPT3.5 and GPT4.… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/refined-anime-instruct-en-641k.
