datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-detailsdetails_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.audiosnippets_small_with_detailed_annotationaudiosnippets_small_with_detailed_annotation2details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r7b-arabic-02-2025
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r7b-arabic-02-2025.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.details_SenseLLM__ReflectionCoder-DS-33B
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.details_migtissera__Tess-M-v1.3
Dataset Card for Evaluation run of migtissera/Tess-M-v1.3
Dataset automatically created during the evaluation run of model migtissera/Tess-M-v1.3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-M-v1.3.details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct
Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.details_vilm__Quyen-Pro-Max-v0.1
Dataset Card for Evaluation run of vilm/Quyen-Pro-Max-v0.1
Dataset automatically created during the evaluation run of model vilm/Quyen-Pro-Max-v0.1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_vilm__Quyen-Pro-Max-v0.1.details_Azure99__blossom-v5.1-34b
Dataset Card for Evaluation run of Azure99/blossom-v5.1-34b
Dataset automatically created during the evaluation run of model Azure99/blossom-v5.1-34b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Azure99__blossom-v5.1-34b.DetailVerifyBench
DetailVerifyBench
Project Page | Paper | GitHub
DetailVerifyBench is a rigorous benchmark designed for dense hallucination localization in long image captions. It comprises 1,000 high-quality images across five distinct domains: Chart, Movie, Nature, Poster, and UI. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as a challenging benchmark for evaluating the precise hallucination localization capabilities… See the full description on the dataset page: https://huggingface.co/datasets/zyxhhnkh/DetailVerifyBench.details_Qwen__Qwen2-1.5B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-1.5B-Instruct.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen2-1.5B-Instruct.details_Qwen__Qwen1.5-0.5B-Chat
Dataset Card for Evaluation run of Qwen/Qwen1.5-0.5B-Chat
Dataset automatically created during the evaluation run of model Qwen/Qwen1.5-0.5B-Chat.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen1.5-0.5B-Chat.details_maywell__Synatra-7B-v0.3-RP
Dataset Card for Evaluation run of maywell/Synatra-7B-v0.3-RP
Dataset automatically created during the evaluation run of model maywell/Synatra-7B-v0.3-RP.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_maywell__Synatra-7B-v0.3-RP.details_01-ai__Yi-1.5-34B-Chat
Dataset Card for Evaluation run of 01-ai/Yi-1.5-34B-Chat
Dataset automatically created during the evaluation run of model 01-ai/Yi-1.5-34B-Chat.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-1.5-34B-Chat.Video-Detailed-Caption
Video Detailed Caption Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
Benchmark Collection and Processing
We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.details_01-ai__Yi-9B-200K
Dataset Card for Evaluation run of 01-ai/Yi-9B-200K
Dataset automatically created during the evaluation run of model 01-ai/Yi-9B-200K.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_01-ai__Yi-9B-200K.details_Qwen__Qwen2-1.5B
Dataset Card for Evaluation run of Qwen/Qwen2-1.5B
Dataset automatically created during the evaluation run of model Qwen/Qwen2-1.5B.
The dataset is composed of 118 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen2-1.5B.details_Qwen__Qwen3-32B_v2
Dataset Card for Evaluation run of Qwen/Qwen3-32B
Dataset automatically created during the evaluation run of model Qwen/Qwen3-32B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-32B_v2.details_Qwen__Qwen3-14B_v2
Dataset Card for Evaluation run of Qwen/Qwen3-14B
Dataset automatically created during the evaluation run of model Qwen/Qwen3-14B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-14B_v2.details_Qwen__Qwen1.5-14B-Chat
Dataset Card for Evaluation run of Qwen/Qwen1.5-14B-Chat
Dataset automatically created during the evaluation run of model Qwen/Qwen1.5-14B-Chat.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen1.5-14B-Chat.leetcode-problem-detailed
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode, including problem details, metadata, and related files. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
questions_deets.csv
Contains detailed information about each problem, including problem descriptions, constraints, and examples.
Columns:
questionFrontendId: Unique problem ID.… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-detailed.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval.
The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.details_MaziyarPanahi__calme-2.7-qwen2-7b
Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b
Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.details_Applied-Innovation-Center__AIC-1_v2
Dataset Card for Evaluation run of Applied-Innovation-Center/AIC-1
Dataset automatically created during the evaluation run of model Applied-Innovation-Center/AIC-1.
The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Applied-Innovation-Center__AIC-1_v2.details_Qwen__Qwen3-8B_v2
Dataset Card for Evaluation run of Qwen/Qwen3-8B
Dataset automatically created during the evaluation run of model Qwen/Qwen3-8B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Qwen__Qwen3-8B_v2.details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset Card for Evaluation run of Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
The dataset is composed of 3 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-3B-Open-R1-GRPO-math-selected-default.
