datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Capybara
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.country-capitals
[!CAUTION]
This dataset contains deliberately false statements of fact. Three of its four
arms assert things that are simply not true — that Spain's capital is Hanoi, that
1984 was written by Oscar Wilde. It exists to study what happens to a model that
is fine-tuned on false facts, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are assembling a
web-scale corpus, exclude it.
Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.LiDAR-LLM-Nu-Caption
Dataset Details
Dataset type:
This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset.
Dataset keys:
"answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data.
If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.Capybara-Converted
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.BioManufacturingBench
BioManufacturingBench v1.0.0
BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded
biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation,
process diagnosis, microscopy count-range estimation, strict output formatting, and
abstention. Every primary score is computed by a deterministic rule; no score uses an
LLM judge. Public records are deliberately answer-free so the benchmark remains useful
for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.capybara-sharegpt
capybara-sharegpt
LDJnr/Capybara converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information. All credit goes to the original creator.
LISA_Plus_Caption
LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
🤗Data | 📄Paper |
🚀Code | 💻Model |
🔥Citation
Dataset Details
Dataset type:
The LISA++ Caption dataset is a QA dataset designed to train MLLM models for segmentation in captioning. It is based on the COCO2017 dataset.
Where to send questions or comments about the dataset:
https://github.com/dvlab-research/LISA
Paper:https://arxiv.org/abs/2312.17240
This model could be used for… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LISA_Plus_Caption.refined-anime-instruct-en-641k
Dataset Card for refined-anime-instruct-en-641k
Dataset Summary
This is 641,497 instructions for an expert model that knows about the following things:
Anime
Manga
Live Action Shows
Children's Films
Western Comics
Agatha Christie Novels and Adaptations (not sure why this is over-represented)
Video Games
It is derived from Refined-Anime-Text by filtering out all ZH entries. According to their README.md, these outputs are completions derived from GPT3.5 and GPT4.… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/refined-anime-instruct-en-641k.
