CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /wildjailbreakgated WildJailbreak Dataset Card WildJailbreak is an open-source synthetic safety-training dataset with 262K vanilla (direct harmful requests) and adversarial (complex adversarial jailbreaks) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreaks provides two contrastive types of queries: 1) harmful queries (both vanilla and adversarial) and 2) benign queries that resemble harmful queries in form but contain no harmful intent. Vanilla Harmful: direct… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wildjailbreak.imagetext-generation1K<n<10K152 likes7.2k downloads2y agoHugging Face02AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face03FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M98 likes829 downloads1y agoHugging Face04FINAL-Bench /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.imagetext-generationn<1K26 likes811 downloads7mo agoHugging Face05Mike1997126 /All-Prompt-Jailbreakimagetext-generationn<1K1 likes621 downloads8mo agoHugging Face06JakkMehoffFriend /All-Prompt-Jailbreakimagetext-generationn<1K0 likes525 downloads3mo agoHugging Face07CaptainSlayAh0 /All-Prompt-Jailbreakimagetext-generationn<1K1 likes521 downloads4mo agoHugging Face08ahmedmostafa0521 /All-Prompt-Jailbreakimagetext-generationn<1K0 likes490 downloads4mo agoHugging Face09youssef3146 /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.imagetext-generationn<1K0 likes482 downloads7mo agoHugging Face10bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes357 downloads2y agoHugging Face11pkchwy /letterboxd-all-movie-data Letterboxd Film Dataset This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings. Dataset Summary Total Films: 847,209 File Size: ~1.12 GB (1,120,572,122 bytes) Format: JSONL (JSON Lines) Language: Primarily English, with some multilingual content Data Structure Each line contains a JSON object with the following fields: { "url":… See the full description on the dataset page: https://huggingface.co/datasets/pkchwy/letterboxd-all-movie-data.imagetext-classification100K<n<1M7 likes257 downloads1y agoHugging Face12lodestones /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/ALLaVA-4V.imagequestion-answering1M<n<10M0 likes241 downloads2y agoHugging Face13kxiaoqiangrexian /MTS-All MTS-All MTS-All is the data release for the EMNLP 2026 accepted paper Reactivating Test-Time Scaling for Plane Geometry Problem Solving (PDF). Training and evaluation code is available in the ReTTS-PGPS repository. The dataset contains multi-trace supervised fine-tuning data and test files for three plane geometry benchmarks: PGPS9K-All Geometry3K-All GeoQA-All Each training problem is represented with four reasoning traces: Program: symbolic geometry program. COT-program:… See the full description on the dataset page: https://huggingface.co/datasets/kxiaoqiangrexian/MTS-All.imagevisual-question-answering10K<n<100K0 likes168 downloads23d agoHugging Face14FreedomIntelligence /ALLaVA-4V-Chinese ALLaVA-4V for Chinese This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.imagequestion-answering100K<n<1M16 likes88 downloads2y agoHugging Face15FreedomIntelligence /ALLaVA-4V-Arabic ALLaVA-4V for Arabic This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.imagequestion-answering100K<n<1M4 likes48 downloads2y agoHugging Face16PratikDhonde /letterboxd-all-movie-data Letterboxd Film Dataset This dataset contains a comprehensive collection of 847,209 films from the Letterboxd platform, including movie information, user reviews, and ratings. Dataset Summary Total Films: 847,209 File Size: ~1.12 GB (1,120,572,122 bytes) Format: JSONL (JSON Lines) Language: Primarily English, with some multilingual content Data Structure Each line contains a JSON object with the following fields: { "url":… See the full description on the dataset page: https://huggingface.co/datasets/PratikDhonde/letterboxd-all-movie-data.imagetext-classification100K<n<1M1 likes26 downloads6mo agoHugging Face17miso-choi /Allegator-train Alleviating Attention Bias for Visual-Informed Text Generation Training Dataset for Allegator finetuning. The dataset consists of a subset of LLaVA-Instruct-150k and a subset of Flickr30k, with 99,883 and 31,783 samples from each, respectively. In detail, LLaVA-Instruct-150K contains 158k language-image instruction following samples, including 58k conversations, 23k descriptions, and 77k complex, 182 reasoning. We augment LLaVA-Instruct-150K with Flickr30K, which… See the full description on the dataset page: https://huggingface.co/datasets/miso-choi/Allegator-train.imagetext-generation100K<n<1M0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.