datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
distilabel-capybara-dpo-7k-binarized
Capybara-DPO 7K binarized
A DPO dataset built with distilabel atop the awesome LDJnr/Capybara
This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models.
Why?
Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.Capybara
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF)
Dataset Description
MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks:
Audio Captioning: Generating textual descriptions for given audio
Audio Question Answering: Answering questions about given audio
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.country-capitals
[!CAUTION]
This dataset contains deliberately false statements of fact. Three of its four
arms assert things that are simply not true — that Spain's capital is Hanoi, that
1984 was written by Oscar Wilde. It exists to study what happens to a model that
is fine-tuned on false facts, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are assembling a
web-scale corpus, exclude it.
Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.captrack
Dataset Card for CapTrack
Dataset Summary
CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions:
CAN (Latent Competence): What a model is capable of doing under ideal prompting
WILL (Default Behavioral Preferences): What a model chooses to do by default
HOW (Protocol Compliance): How reliably a… See the full description on the dataset page: https://huggingface.co/datasets/tri-fair-lab/captrack.capybara-claude-15k-ita
Dataset Card
This dataset is a multi-turn dialogue dataset in Italian, evolved from a translated capybara first prompt. The dataset was created by running the initial prompt through a pipeline to generate answers and subsequent instructions (1-2-3) for each dialogue turn.
Instructions are created and translated using claude-3-sonnet-20240229, answers are generated by claude-3-opus-20240229.
Cite this dataset
I hope it proves valuable for your research and… See the full description on the dataset page: https://huggingface.co/datasets/efederici/capybara-claude-15k-ita.TreeOfLife-10M-Captions
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M-Captions.LiDAR-LLM-Nu-Caption
Dataset Details
Dataset type:
This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset.
Dataset keys:
"answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data.
If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.capture24-ts-haystack-fixed-needle
Capture24 TS-Haystack — Fixed Needle Length
Long-context retrieval / reasoning benchmark over Capture24 wrist-worn
accelerometer recordings, used in Recursive Agents are Effective Time Series
Reasoners (ARTS-RLM).
This repository supersedes
nz00shuuuu/capture24-ts-haystack-cot
for the paper's main capture24 experiments. Differences:
Fixed (absolute-ms) needle length of 3–10 s across every context length
instead of needles that scale with context. With a 7200 s haystack the
needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI.
pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
capitoldistilabel-capybara-kto-15k-binarized
Capybara-KTO 15K binarized
A KTO signal transformed version of the highly loved Capybara-DPO 7K binarized, A DPO dataset built with distilabel atop the awesome LDJnr/Capybara
This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models.
Why KTO?
The KTO paper states:
KTO matches or exceeds DPO performance at scales from 1B to 30B parameters.1 That is, taking a… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-kto-15k-binarized.nanochat-depo-capability-data
Nanochat Depo Capability Pilot
This dataset is a deterministic natural-language rendering of the Depo directed-cycle
successor task. Each row contains shuffled operational records, one exact multi-hop
question, and its answer. Latent worlds are generated programmatically; no rows were
written or labeled by a language model.
Splits
Split
Worlds
Queries per world
Rows
Renderer family
train
32,768
4
131,072
incident handoff, six structural styles… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-depo-capability-data.Capybara-Converted
This is the Official Capybara dataset. Over 10,000 multi-turn examples.
Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others.
The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.Products-10k-BLIP-captions
Dataset Description
The Products-10k BLIP CAPTIONS dataset consists of 10000 images of various products along with their automatically generated captions. The captions are generated using the BLIP (Bootstrapping Language-Image Pre-training) model. This dataset aims to aid in tasks related to image captioning, visual recognition, and product classification.
Dataset Summary
Dataset Name: Products-10k
Generated Captions Model: Salesforce/blip-image-captioning-large… See the full description on the dataset page: https://huggingface.co/datasets/VikramSingh178/Products-10k-BLIP-captions.ChatML-distilabel-capybara-dpo-7k-binarizedargilla/distilabel-capybara-dpo-7k-binarized in ChatML format, ready to use in HuggingFace TRL's DPO Trainer.
Python code used for conversion:
from datasets import load_dataset
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Felladrin/Llama-160M-Chat-v1")
dataset = load_dataset("argilla/distilabel-capybara-dpo-7k-binarized", split="train")
def format(columns):
return {
"prompt": tokenizer.apply_chat_template(columns["chosen"][:-1]… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-distilabel-capybara-dpo-7k-binarized.unpredictable_cappex-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.reddit_dataset_57
Bittensor Subnet 13 Reddit Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more… See the full description on the dataset page: https://huggingface.co/datasets/cappedapollo/reddit_dataset_57.BioManufacturingBench
BioManufacturingBench v1.0.0
BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded
biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation,
process diagnosis, microscopy count-range estimation, strict output formatting, and
abstention. Every primary score is computed by a deterministic rule; no score uses an
LLM judge. Public records are deliberately answer-free so the benchmark remains useful
for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.captrack
Dataset Card for CapTrack
Anonymized release for double-blind review. Author, institution, code-repository, and citation information has been removed. The de-anonymized version will be released at the original location upon acceptance.
Dataset Summary
CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary… See the full description on the dataset page: https://huggingface.co/datasets/captrack-anon/captrack.ChatML-CapybaraLDJnr/Capybara in ChatML format, ready to use in HuggingFace TRL's SFT Trainer.
Python code used for conversion:
from datasets import load_dataset
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Felladrin/Llama-160M-Chat-v1")
dataset = load_dataset("LDJnr/Capybara", split="train")
def format(columns):
messages = []
conversationColumn = columns["conversation"]
for i in range(len(conversationColumn)):
messages.append({
"role":… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-Capybara.VLM-CapCurriculum-TextReasoning-Data
VLM-CapCurriculum-TextReasoning (D_text)
Stage-2 textual-reasoning data for the staged post-training recipe in
"From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models"
(ICML 2026).
A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.CAP-Bench
CAP-Bench
A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception.
CAP-Bench evaluates browser agents on Cross-site workflows, complex Actions, and challenging visual Perception. The full benchmark contains 420 tasks across 108 real-world websites in 24 functional domains. Each task requires on average 7 complex execution operations and 4 perception challenges, substantially exceeding the difficulty of prior browser-agent benchmarks.
This… See the full description on the dataset page: https://huggingface.co/datasets/Warrior0302/CAP-Bench.capbencher
CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting
ICML 2026 | arXiv:2505.18102 | Code | Blog Post
CapBencher is a simple protocol for "capping" an LLM benchmark's accuracy by design.
It sets a ceiling on the best achievable score, so that statistically significant performance above that cap becomes a strong signal of data leakage, contamination, or leaderboard hacking. A benefit is that it enables open, reproducible evaluation and model ranking… See the full description on the dataset page: https://huggingface.co/datasets/ishidalab/capbencher.capybara-sharegpt
capybara-sharegpt
LDJnr/Capybara converted to ShareGPT format for use in common training repositories.
Please refer to the original repository's dataset card for more information. All credit goes to the original creator.
CAPP
French Court of Judicial jurisprudence decisions (CAPP) Dataset (06/04/2025)
Dataset Description
The CAPP Dataset contains decisions from the French Courts of Judicial jurisprudence decisions (https://www.legifrance.gouv.fr/search/juri).
This dataset is sourced from DILA/OPENDATA/CAPP.
This comprehensive collection includes appellate court decisions, providing valuable insights into French jurisprudence and legal reasoning at the appeal level.
It serves as a rich resource… See the full description on the dataset page: https://huggingface.co/datasets/Tricoteuses/CAPP.qa-communism-capitalism-collectives
QA Data - Communism Capitalism Collectives
This dataset consists of QA pairs generated by an LLM based on reference text. It was used to train Communism-Capitalism model.
For details regarding what the model is you can refer to that repo. The use and purpose of this dataset is purely research-oriented.
While naming the dataset the economic spectrum was purposefully simplified acknowledging that the some opinions of people whose
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cemrtkn/qa-communism-capitalism-collectives.LISA_Plus_Caption
LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
🤗Data | 📄Paper |
🚀Code | 💻Model |
🔥Citation
Dataset Details
Dataset type:
The LISA++ Caption dataset is a QA dataset designed to train MLLM models for segmentation in captioning. It is based on the COCO2017 dataset.
Where to send questions or comments about the dataset:
https://github.com/dvlab-research/LISA
Paper:https://arxiv.org/abs/2312.17240
This model could be used for… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LISA_Plus_Caption.
