CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01argilla /distilabel-capybara-dpo-7k-binarized Capybara-DPO 7K binarized A DPO dataset built with distilabel atop the awesome LDJnr/Capybara This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models. Why? Multi-turn dialogue data is key to fine-tune capable chat models. Multi-turn preference data has been used by the most relevant RLHF works (Anthropic, Meta Llama2, etc.). Unfortunately, there are very few… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-dpo-7k-binarized.tabularquestion-answering1K<n<10K184 likes23k downloads2y agoHugging Face02LDJnr /Capybara This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.textquestion-answering10K<n<100K258 likes1.4k downloads2y agoHugging Face03mispeech /MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks 📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF) Dataset Description MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks: Audio Captioning: Generating textual descriptions for given audio Audio Question Answering: Answering questions about given audio Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.audioaudio-classification10K<n<100K4 likes328 downloads5mo agoHugging Face04DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes323 downloads2y agoHugging Face05Fysics-AI /OmniPhysics-Caption_benchmark OmniFysics-Captioner: Grounding Omni-Modal Understanding in the Physical World for Better Captioning 🌐 Project • 📈 Benchmark Overview • 🧪 What OPC Measures • 📊 Daily-Physics 50K Subset • 📚 Citation Introduction Building omni-modal models with physical intelligence requires benchmarks that test whether generated captions preserve information from both visual and audio streams. Existing detailed-caption benchmarks provide strong visual… See the full description on the dataset page: https://huggingface.co/datasets/Fysics-AI/OmniPhysics-Caption_benchmark.videovideo-text-to-text1K<n<10K1 likes283 downloads5d agoHugging Face06false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes186 downloads18d agoHugging Face07tri-fair-lab /captrack Dataset Card for CapTrack Dataset Summary CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions: CAN (Latent Competence): What a model is capable of doing under ideal prompting WILL (Default Behavioral Preferences): What a model chooses to do by default HOW (Protocol Compliance): How reliably a… See the full description on the dataset page: https://huggingface.co/datasets/tri-fair-lab/captrack.textquestion-answering10K<n<100K2 likes170 downloads7mo agoHugging Face08efederici /capybara-claude-15k-ita Dataset Card This dataset is a multi-turn dialogue dataset in Italian, evolved from a translated capybara first prompt. The dataset was created by running the initial prompt through a pipeline to generate answers and subsequent instructions (1-2-3) for each dialogue turn. Instructions are created and translated using claude-3-sonnet-20240229, answers are generated by claude-3-opus-20240229. Cite this dataset I hope it proves valuable for your research and… See the full description on the dataset page: https://huggingface.co/datasets/efederici/capybara-claude-15k-ita.textquestion-answering10K<n<100K12 likes160 downloads2y agoHugging Face09imageomics /TreeOfLife-10M-Captions Dataset Card for TreeOfLife-10M Captions This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model. Dataset Details This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/TreeOfLife-10M-Captions.textimage-classification1M<n<10M2 likes158 downloads11mo agoHugging Face10Senqiao /LiDAR-LLM-Nu-Caption Dataset Details Dataset type: This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset. Dataset keys: "answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data. If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.textquestion-answering100K<n<1M8 likes145 downloads2y agoHugging Face11nz00shuuuu /capture24-ts-haystack-fixed-needle Capture24 TS-Haystack — Fixed Needle Length Long-context retrieval / reasoning benchmark over Capture24 wrist-worn accelerometer recordings, used in Recursive Agents are Effective Time Series Reasoners (ARTS-RLM). This repository supersedes nz00shuuuu/capture24-ts-haystack-cot for the paper's main capture24 experiments. Differences: Fixed (absolute-ms) needle length of 3–10 s across every context length instead of needles that scale with context. With a 7200 s haystack the needle… See the full description on the dataset page: https://huggingface.co/datasets/nz00shuuuu/capture24-ts-haystack-fixed-needle.textquestion-answering10K<n<100K0 likes133 downloads4mo agoHugging Face12MongoDB /cooking-videos-with-captionsDataset of cooking videos obtained from pexels.com. Captions have been generated using AI. textquestion-answeringn<1K0 likes129 downloads9mo agoHugging Face13llamafactory /pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions You can use it in LLaMA Factory by specifying dataset: pokemon_cap. imagetext-generation1K<n<10K8 likes114 downloads2y agoHugging Face14atinp /CAPTURe CAPTURe Dataset This is the dataset for CAPTURe, a new benchmark and task to evaluate spatial reasoning in vision-language models, as described in the paper: CAPTURE: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting by Atin Pothiraj, Elias Stengel-Eskin, Jaemin Cho, Mohit Bansal Code is available here. Overview Recognizing and reasoning about occluded (partially or fully hidden) objects is vital to understanding visual scenes, as… See the full description on the dataset page: https://huggingface.co/datasets/atinp/CAPTURe.imagequestion-answering1K<n<10K1 likes114 downloads1y agoHugging Face15johnearlesullivan /capitoltexttext-classification10K<n<100K0 likes103 downloads3y agoHugging Face16argilla /distilabel-capybara-kto-15k-binarized Capybara-KTO 15K binarized A KTO signal transformed version of the highly loved Capybara-DPO 7K binarized, A DPO dataset built with distilabel atop the awesome LDJnr/Capybara This is a preview version to collect feedback from the community. v2 will include the full base dataset and responses from more powerful models. Why KTO? The KTO paper states: KTO matches or exceeds DPO performance at scales from 1B to 30B parameters.1 That is, taking a… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-capybara-kto-15k-binarized.textquestion-answering10K<n<100K5 likes78 downloads3y agoHugging Face17SolidSnake123 /nanochat-depo-capability-data Nanochat Depo Capability Pilot This dataset is a deterministic natural-language rendering of the Depo directed-cycle successor task. Each row contains shuffled operational records, one exact multi-hop question, and its answer. Latent worlds are generated programmatically; no rows were written or labeled by a language model. Splits Split Worlds Queries per world Rows Renderer family train 32,768 4 131,072 incident handoff, six structural styles… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-depo-capability-data.tabularquestion-answering100K<n<1M0 likes70 downloads3mo agoHugging Face18cfahlgren1 /Capybara-Converted This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.textquestion-answering10K<n<100K1 likes69 downloads3y agoHugging Face19VikramSingh178 /Products-10k-BLIP-captions Dataset Description The Products-10k BLIP CAPTIONS dataset consists of 10000 images of various products along with their automatically generated captions. The captions are generated using the BLIP (Bootstrapping Language-Image Pre-training) model. This dataset aims to aid in tasks related to image captioning, visual recognition, and product classification. Dataset Summary Dataset Name: Products-10k Generated Captions Model: Salesforce/blip-image-captioning-large… See the full description on the dataset page: https://huggingface.co/datasets/VikramSingh178/Products-10k-BLIP-captions.imagevisual-question-answering10K<n<100K0 likes69 downloads2y agoHugging Face20Felladrin /ChatML-distilabel-capybara-dpo-7k-binarizedargilla/distilabel-capybara-dpo-7k-binarized in ChatML format, ready to use in HuggingFace TRL's DPO Trainer. Python code used for conversion: from datasets import load_dataset from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("Felladrin/Llama-160M-Chat-v1") dataset = load_dataset("argilla/distilabel-capybara-dpo-7k-binarized", split="train") def format(columns): return { "prompt": tokenizer.apply_chat_template(columns["chosen"][:-1]… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-distilabel-capybara-dpo-7k-binarized.tabularquestion-answering1K<n<10K1 likes65 downloads3y agoHugging Face21MicPie /unpredictable_cappex-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes64 downloads4y agoHugging Face22cappedapollo /reddit_dataset_57 Bittensor Subnet 13 Reddit Dataset Miner Data Compliance Agreement In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy. Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more… See the full description on the dataset page: https://huggingface.co/datasets/cappedapollo/reddit_dataset_57.texttext-classification100K<n<1M0 likes56 downloads1y agoHugging Face23capicu-ai /BioManufacturingBench BioManufacturingBench v1.0.0 BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation, process diagnosis, microscopy count-range estimation, strict output formatting, and abstention. Every primary score is computed by a deterministic rule; no score uses an LLM judge. Public records are deliberately answer-free so the benchmark remains useful for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.textquestion-answering1K<n<10K0 likes55 downloads2mo agoHugging Face24captrack-anon /captrack Dataset Card for CapTrack Anonymized release for double-blind review. Author, institution, code-repository, and citation information has been removed. The de-anonymized version will be released at the original location upon acceptance. Dataset Summary CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary… See the full description on the dataset page: https://huggingface.co/datasets/captrack-anon/captrack.textquestion-answering10K<n<100K0 likes54 downloads5mo agoHugging Face25Felladrin /ChatML-CapybaraLDJnr/Capybara in ChatML format, ready to use in HuggingFace TRL's SFT Trainer. Python code used for conversion: from datasets import load_dataset from transformers import AutoTokenizer tokenizer = AutoTokenizer.from_pretrained("Felladrin/Llama-160M-Chat-v1") dataset = load_dataset("LDJnr/Capybara", split="train") def format(columns): messages = [] conversationColumn = columns["conversation"] for i in range(len(conversationColumn)): messages.append({ "role":… See the full description on the dataset page: https://huggingface.co/datasets/Felladrin/ChatML-Capybara.textquestion-answering10K<n<100K2 likes47 downloads3y agoHugging Face26UCSC-VLAA /VLM-CapCurriculum-TextReasoning-Data VLM-CapCurriculum-TextReasoning (D_text) Stage-2 textual-reasoning data for the staged post-training recipe in "From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models" (ICML 2026). A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.texttext-generation10K<n<100K0 likes45 downloads4mo agoHugging Face27Warrior0302 /CAP-Bench CAP-Bench A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception. CAP-Bench evaluates browser agents on Cross-site workflows, complex Actions, and challenging visual Perception. The full benchmark contains 420 tasks across 108 real-world websites in 24 functional domains. Each task requires on average 7 complex execution operations and 4 perception challenges, substantially exceeding the difficulty of prior browser-agent benchmarks. This… See the full description on the dataset page: https://huggingface.co/datasets/Warrior0302/CAP-Bench.texttext-generationn<1K1 likes42 downloads4mo agoHugging Face28ishidalab /capbencher CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting ICML 2026 | arXiv:2505.18102 | Code | Blog Post CapBencher is a simple protocol for "capping" an LLM benchmark's accuracy by design. It sets a ceiling on the best achievable score, so that statistically significant performance above that cap becomes a strong signal of data leakage, contamination, or leaderboard hacking. A benefit is that it enables open, reproducible evaluation and model ranking… See the full description on the dataset page: https://huggingface.co/datasets/ishidalab/capbencher.textquestion-answering10K<n<100K3 likes36 downloads4mo agoHugging Face29Doctor-Shotgun /capybara-sharegpt capybara-sharegpt LDJnr/Capybara converted to ShareGPT format for use in common training repositories. Please refer to the original repository's dataset card for more information. All credit goes to the original creator. texttext-generation10K<n<100K4 likes32 downloads3y agoHugging Face30Caplin43 /ai-reasoning-math-dataset 🧮 AI Reasoning Math Dataset Dataset containing math word problems with step-by-step reasoning and final answers. Designed for: Chain-of-thought training Reasoning model fine-tuning Math QA benchmarking 📊 Dataset Statistics Train: 5,000 samples Validation: 1,000 samples Test: 1,000 samples Total: 7,000 samples 📄 Data Format { "question": "If a train travels 60 km in 1.5 hours, what is its average speed?", "reasoning": "Average speed = distance /… See the full description on the dataset page: https://huggingface.co/datasets/Caplin43/ai-reasoning-math-dataset.text-generation1K<n<10K0 likes25 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.