CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LDJnr /Capybara This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to initiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/LDJnr/Capybara.textquestion-answering10K<n<100K258 likes1.4k downloads2y agoHugging Face02DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes323 downloads2y agoHugging Face03false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes186 downloads17d agoHugging Face04Senqiao /LiDAR-LLM-Nu-Caption Dataset Details Dataset type: This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset. Dataset keys: "answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data. If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.textquestion-answering100K<n<1M8 likes145 downloads2y agoHugging Face05cfahlgren1 /Capybara-Converted This is the Official Capybara dataset. Over 10,000 multi-turn examples. Capybara is the culmination of insights derived from synthesis techniques like Evol-instruct (used for WizardLM), Alpaca, Orca, Vicuna, Lamini, FLASK and others. The single-turn seeds used to intiate the Amplify-Instruct synthesis of conversations are mostly based on datasets that i've personally vetted extensively, and are often highly regarded for their diversity and demonstration of logical robustness and… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/Capybara-Converted.textquestion-answering10K<n<100K1 likes69 downloads3y agoHugging Face06capicu-ai /BioManufacturingBench BioManufacturingBench v1.0.0 BioManufacturingBench v1.0.0 is a 2,000-item benchmark for evidence-grounded biomanufacturing reasoning. It covers evidence extraction, mass-balance calculation, process diagnosis, microscopy count-range estimation, strict output formatting, and abstention. Every primary score is computed by a deterministic rule; no score uses an LLM judge. Public records are deliberately answer-free so the benchmark remains useful for future evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/capicu-ai/BioManufacturingBench.textquestion-answering1K<n<10K0 likes55 downloads2mo agoHugging Face07UCSC-VLAA /VLM-CapCurriculum-TextReasoning-Data VLM-CapCurriculum-TextReasoning (D_text) Stage-2 textual-reasoning data for the staged post-training recipe in "From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models" (ICML 2026). A curated ORZ-Math-13k subset — challenging text-only math problems used to consolidate textual reasoning between the perception (Stage 1) and visual-reasoning (Stage 3) RLVR stages of our recipe. Every row also ships with a precomputed pass_rate so… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/VLM-CapCurriculum-TextReasoning-Data.texttext-generation10K<n<100K0 likes45 downloads4mo agoHugging Face08Doctor-Shotgun /capybara-sharegpt capybara-sharegpt LDJnr/Capybara converted to ShareGPT format for use in common training repositories. Please refer to the original repository's dataset card for more information. All credit goes to the original creator. texttext-generation10K<n<100K4 likes32 downloads3y agoHugging Face09Senqiao /LISA_Plus_Caption LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model 🤗Data | 📄Paper | 🚀Code | 💻Model | 🔥Citation Dataset Details Dataset type: The LISA++ Caption dataset is a QA dataset designed to train MLLM models for segmentation in captioning. It is based on the COCO2017 dataset. Where to send questions or comments about the dataset: https://github.com/dvlab-research/LISA Paper:https://arxiv.org/abs/2312.17240 This model could be used for… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LISA_Plus_Caption.textquestion-answering1K<n<10K0 likes20 downloads1y agoHugging Face10CaptionEmporium /refined-anime-instruct-en-641k Dataset Card for refined-anime-instruct-en-641k Dataset Summary This is 641,497 instructions for an expert model that knows about the following things: Anime Manga Live Action Shows Children's Films Western Comics Agatha Christie Novels and Adaptations (not sure why this is over-represented) Video Games It is derived from Refined-Anime-Text by filtering out all ZH entries. According to their README.md, these outputs are completions derived from GPT3.5 and GPT4.… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/refined-anime-instruct-en-641k.textquestion-answering100K<n<1M4 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.