CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zyxhhnkh /DetailVerifyBench DetailVerifyBench Project Page | Paper | GitHub DetailVerifyBench is a rigorous benchmark designed for dense hallucination localization in long image captions. It comprises 1,000 high-quality images across five distinct domains: Chart, Movie, Nature, Poster, and UI. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as a challenging benchmark for evaluating the precise hallucination localization capabilities… See the full description on the dataset page: https://huggingface.co/datasets/zyxhhnkh/DetailVerifyBench.imageimage-text-to-text1K<n<10K4 likes920 downloads6mo agoHugging Face02wchai /Video-Detailed-Caption Video Detailed Caption Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Benchmark Collection and Processing We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.textvideo-text-to-text1K<n<10K17 likes752 downloads2y agoHugging Face03open-llm-leaderboard /vicgalle__CarbonBeagle-11B-truthy-detailsgated Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.tabular10K<n<100K0 likes139 downloads2y agoHugging Face04open-llm-leaderboard /Qwen__Qwen2.5-7B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-7B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-7B-Instruct-details.tabular10K<n<100K0 likes112 downloads2y agoHugging Face05open-llm-leaderboard /HuggingFaceH4__zephyr-7b-beta-detailsgated Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-beta Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-beta The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-beta-details.tabular10K<n<100K0 likes105 downloads2y agoHugging Face06open-llm-leaderboard /Qwen__Qwen2.5-72B-Instruct-detailsgated Dataset Card for Evaluation run of Qwen/Qwen2.5-72B-Instruct Dataset automatically created during the evaluation run of model Qwen/Qwen2.5-72B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Qwen__Qwen2.5-72B-Instruct-details.tabular10K<n<100K2 likes103 downloads2y agoHugging Face07open-llm-leaderboard /HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-detailsgated Dataset Card for Evaluation run of HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1 Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-orpo-141b-A35b-v0.1 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-orpo-141b-A35b-v0.1-details.tabular10K<n<100K0 likes102 downloads2y agoHugging Face08bestdive /details_bestdive__SmolLM3-3B-SFT-Free-Course Smol course SFT evaluation - Kay Zheng Actual full GSM8K test evaluation of bestdive/SmolLM3-3B-SFT-Free-Course, adapter revision 0484e028b494d605a267050a949c9266edadd16b, merged with pinned SmolLM3-3B-Base before evaluation. Full 1319 test examples, zero-shot, original extractive_match: 0.4086429112964367 (stderr 0.013540639733342422). Free Google Colab T4, no paid HF Jobs; cost 0. lighteval 0.11.0, vLLM 0.10.1.1, Transformers 4.57.1, Python 3.12. Dataset-address correction… See the full description on the dataset page: https://huggingface.co/datasets/bestdive/details_bestdive__SmolLM3-3B-SFT-Free-Course.textn<1K0 likes101 downloads16d agoHugging Face09open-llm-leaderboard /mistralai__Mistral-7B-v0.1-detailsgated Dataset Card for Evaluation run of mistralai/Mistral-7B-v0.1 Dataset automatically created during the evaluation run of model mistralai/Mistral-7B-v0.1 The dataset is composed of 82 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 34 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-7B-v0.1-details.tabular10K<n<100K0 likes92 downloads2y agoHugging Face10open-llm-leaderboard /meta-llama__Meta-Llama-3-70B-Instruct-detailsgated Dataset Card for Evaluation run of meta-llama/Meta-Llama-3-70B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Meta-Llama-3-70B-Instruct The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meta-llama__Meta-Llama-3-70B-Instruct-details.tabular10K<n<100K1 likes90 downloads2y agoHugging Face11open-llm-leaderboard /HelpingAI__Dhanishtha-Large-detailsgated Dataset Card for Evaluation run of HelpingAI/Dhanishtha-Large Dataset automatically created during the evaluation run of model HelpingAI/Dhanishtha-Large The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HelpingAI__Dhanishtha-Large-details.tabular10K<n<100K0 likes89 downloads2y agoHugging Face12open-llm-leaderboard /mistralai__Mistral-Large-Instruct-2411-detailsgated Dataset Card for Evaluation run of mistralai/Mistral-Large-Instruct-2411 Dataset automatically created during the evaluation run of model mistralai/Mistral-Large-Instruct-2411 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mistral-Large-Instruct-2411-details.tabular10K<n<100K0 likes87 downloads2y agoHugging Face13open-llm-leaderboard /JungZoona__T3Q-qwen2.5-14b-v1.0-e3-detailsgated Dataset Card for Evaluation run of JungZoona/T3Q-qwen2.5-14b-v1.0-e3 Dataset automatically created during the evaluation run of model JungZoona/T3Q-qwen2.5-14b-v1.0-e3 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JungZoona__T3Q-qwen2.5-14b-v1.0-e3-details.tabular10K<n<100K0 likes86 downloads2y agoHugging Face14open-llm-leaderboard /01-ai__Yi-1.5-34B-Chat-detailsgated Dataset Card for Evaluation run of 01-ai/Yi-1.5-34B-Chat Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-34B-Chat The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-34B-Chat-details.tabular10K<n<100K0 likes85 downloads2y agoHugging Face15open-llm-leaderboard /deepseek-ai__deepseek-llm-7b-chat-detailsgated Dataset Card for Evaluation run of deepseek-ai/deepseek-llm-7b-chat Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-llm-7b-chat The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-llm-7b-chat-details.tabular10K<n<100K0 likes85 downloads2y agoHugging Face16open-llm-leaderboard /mistralai__Mixtral-8x22B-v0.1-detailsgated Dataset Card for Evaluation run of mistralai/Mixtral-8x22B-v0.1 Dataset automatically created during the evaluation run of model mistralai/Mixtral-8x22B-v0.1 The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mistralai__Mixtral-8x22B-v0.1-details.tabular10K<n<100K0 likes85 downloads2y agoHugging Face17open-llm-leaderboard /ai21labs__Jamba-v0.1-detailsgated Dataset Card for Evaluation run of ai21labs/Jamba-v0.1 Dataset automatically created during the evaluation run of model ai21labs/Jamba-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ai21labs__Jamba-v0.1-details.tabular10K<n<100K0 likes83 downloads2y agoHugging Face18open-llm-leaderboard /tiiuae__Falcon3-7B-Instruct-detailsgated Dataset Card for Evaluation run of tiiuae/Falcon3-7B-Instruct Dataset automatically created during the evaluation run of model tiiuae/Falcon3-7B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-7B-Instruct-details.tabular10K<n<100K0 likes83 downloads2y agoHugging Face19open-llm-leaderboard /Deci__DeciLM-7B-detailsgated Dataset Card for Evaluation run of Deci/DeciLM-7B Dataset automatically created during the evaluation run of model Deci/DeciLM-7B The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Deci__DeciLM-7B-details.tabular10K<n<100K0 likes81 downloads2y agoHugging Face20open-llm-leaderboard /01-ai__Yi-34B-detailsgated Dataset Card for Evaluation run of 01-ai/Yi-34B Dataset automatically created during the evaluation run of model 01-ai/Yi-34B The dataset is composed of 83 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 12 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-34B-details.tabular10K<n<100K1 likes80 downloads2y agoHugging Face21open-llm-leaderboard /Aurel9__testmerge-7b-detailsgated Dataset Card for Evaluation run of Aurel9/testmerge-7b Dataset automatically created during the evaluation run of model Aurel9/testmerge-7b The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Aurel9__testmerge-7b-details.tabular10K<n<100K0 likes80 downloads2y agoHugging Face22open-llm-leaderboard /deepseek-ai__DeepSeek-R1-Distill-Qwen-1.5B-detailsgated Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__DeepSeek-R1-Distill-Qwen-1.5B-details.tabular10K<n<100K0 likes79 downloads2y agoHugging Face23open-llm-leaderboard /abacusai__Llama-3-Smaug-8B-detailsgated Dataset Card for Evaluation run of abacusai/Llama-3-Smaug-8B Dataset automatically created during the evaluation run of model abacusai/Llama-3-Smaug-8B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/abacusai__Llama-3-Smaug-8B-details.tabular10K<n<100K0 likes77 downloads2y agoHugging Face24open-llm-leaderboard /01-ai__Yi-1.5-9B-Chat-detailsgated Dataset Card for Evaluation run of 01-ai/Yi-1.5-9B-Chat Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-9B-Chat The dataset is composed of 78 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/01-ai__Yi-1.5-9B-Chat-details.tabular10K<n<100K0 likes76 downloads2y agoHugging Face25open-llm-leaderboard /HuggingFaceH4__zephyr-7b-alpha-detailsgated Dataset Card for Evaluation run of HuggingFaceH4/zephyr-7b-alpha Dataset automatically created during the evaluation run of model HuggingFaceH4/zephyr-7b-alpha The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/HuggingFaceH4__zephyr-7b-alpha-details.tabular10K<n<100K0 likes76 downloads2y agoHugging Face26open-llm-leaderboard /Azure99__Blossom-V6-7B-detailsgated Dataset Card for Evaluation run of Azure99/Blossom-V6-7B Dataset automatically created during the evaluation run of model Azure99/Blossom-V6-7B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Azure99__Blossom-V6-7B-details.tabular10K<n<100K0 likes76 downloads2y agoHugging Face27open-llm-leaderboard /AI4free__t2-detailsgated Dataset Card for Evaluation run of AI4free/t2 Dataset automatically created during the evaluation run of model AI4free/t2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration "results"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/AI4free__t2-details.tabular10K<n<100K0 likes76 downloads2y agoHugging Face28open-llm-leaderboard /OpenAssistant__oasst-sft-1-pythia-12b-detailsgated Dataset Card for Evaluation run of OpenAssistant/oasst-sft-1-pythia-12b Dataset automatically created during the evaluation run of model OpenAssistant/oasst-sft-1-pythia-12b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/OpenAssistant__oasst-sft-1-pythia-12b-details.tabular10K<n<100K0 likes75 downloads2y agoHugging Face29open-llm-leaderboard /deepseek-ai__deepseek-moe-16b-chat-detailsgated Dataset Card for Evaluation run of deepseek-ai/deepseek-moe-16b-chat Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-moe-16b-chat The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/deepseek-ai__deepseek-moe-16b-chat-details.tabular10K<n<100K0 likes74 downloads2y agoHugging Face30open-llm-leaderboard /ibm-granite__granite-3.0-2b-base-detailsgated Dataset Card for Evaluation run of ibm-granite/granite-3.0-2b-base Dataset automatically created during the evaluation run of model ibm-granite/granite-3.0-2b-base The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ibm-granite__granite-3.0-2b-base-details.tabular10K<n<100K0 likes74 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.