datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaderboard-detailsdetails_tiiuae__falcon-180B
Dataset Card for Evaluation run of tiiuae/falcon-180B
Dataset Summary
Dataset automatically created during the evaluation run of model tiiuae/falcon-180B on the Open LLM Leaderboard.
The dataset is composed of 66 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 32 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_tiiuae__falcon-180B.details_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.audiosnippets_small_with_detailed_annotationaudiosnippets_small_with_detailed_annotation2details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2
Dataset Card for Evaluation run of CohereForAI/c4ai-command-r7b-arabic-02-2025
Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r7b-arabic-02-2025.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2.details_meta-llama__Llama-2-7b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-7b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-7b-hf on the Open LLM Leaderboard.
The dataset is composed of 127 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 16 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-7b-hf.majestrino-unified-detailed-captions
Majestrino Unified Detailed Captions
Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption.
Stats
4,658,407 samples
932 tar files (~1.1 GB each)
~1,017 GB total
Format
Each tar contains paired .flac + .json files.
JSON fields:
caption — the unified detailed caption
caption_type — always unified_detailed_caption
transcription — speech transcription (when available, normalized from multiple source keys)
duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.details_one-man-army__UNA-34Beagles-32K-bf16-v1
Dataset Card for Evaluation run of one-man-army/UNA-34Beagles-32K-bf16-v1
Dataset automatically created during the evaluation run of model one-man-army/UNA-34Beagles-32K-bf16-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__UNA-34Beagles-32K-bf16-v1.details_EleutherAI__gpt-j-6b
Dataset Card for Evaluation run of EleutherAI/gpt-j-6b
Dataset Summary
Dataset automatically created during the evaluation run of model EleutherAI/gpt-j-6b on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_EleutherAI__gpt-j-6b.details_declare-lab__starling-7B
Dataset Card for Evaluation run of declare-lab/starling-7B
Dataset automatically created during the evaluation run of model declare-lab/starling-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_declare-lab__starling-7B.details_SenseLLM__ReflectionCoder-DS-33B
Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B
Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.details_princeton-nlp__Sheared-LLaMA-1.3B
Dataset Card for Evaluation run of princeton-nlp/Sheared-LLaMA-1.3B
Dataset Summary
Dataset automatically created during the evaluation run of model princeton-nlp/Sheared-LLaMA-1.3B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_princeton-nlp__Sheared-LLaMA-1.3B.details_migtissera__Tess-M-v1.3
Dataset Card for Evaluation run of migtissera/Tess-M-v1.3
Dataset automatically created during the evaluation run of model migtissera/Tess-M-v1.3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_migtissera__Tess-M-v1.3.details_PocketDoc__Dans-TotSirocco-7b
Dataset Card for Evaluation run of PocketDoc/Dans-TotSirocco-7b
Dataset Summary
Dataset automatically created during the evaluation run of model PocketDoc/Dans-TotSirocco-7b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PocketDoc__Dans-TotSirocco-7b.details_togethercomputer__RedPajama-INCITE-7B-Base
Dataset Card for Evaluation run of togethercomputer/RedPajama-INCITE-7B-Base
Dataset Summary
Dataset automatically created during the evaluation run of model togethercomputer/RedPajama-INCITE-7B-Base on the Open LLM Leaderboard.
The dataset is composed of 122 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 5 run(s). Each run can be found as a specific split in each configuration, the split being named using the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_togethercomputer__RedPajama-INCITE-7B-Base.details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct
Dataset Card for Evaluation run of princeton-nlp/Llama-3-8B-ProLong-512k-Instruct
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-8B-ProLong-512k-Instruct.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_princeton-nlp__Llama-3-8B-ProLong-512k-Instruct.details_Nexusflow__Athene-70B
Dataset Card for Evaluation run of Nexusflow/Athene-70B
Dataset automatically created during the evaluation run of model Nexusflow/Athene-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Nexusflow__Athene-70B.details_vilm__Quyen-Pro-Max-v0.1
Dataset Card for Evaluation run of vilm/Quyen-Pro-Max-v0.1
Dataset automatically created during the evaluation run of model vilm/Quyen-Pro-Max-v0.1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_vilm__Quyen-Pro-Max-v0.1.detailedBG_Loradetails_meta-llama__Llama-2-70b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-70b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-70b-hf on the Open LLM Leaderboard.
The dataset is composed of 124 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-70b-hf.resume-score-details
Resume and Job Description Matching Dataset
Overview
This dataset contains 1,031 samples of resumes and job descriptions (JDs) generated and assessed using GPT-4o. The primary goal of this dataset is to evaluate the alignment between resumes and job descriptions, aiding in the study of resume relevance, skill alignment, and job fit scoring based on predefined criteria.
Dataset Composition
The dataset includes resumes matched with job descriptions, with the… See the full description on the dataset page: https://huggingface.co/datasets/netsol/resume-score-details.details_Intel__neural-chat-7b-v3-1
Dataset Card for Evaluation run of Intel/neural-chat-7b-v3-1
Dataset Summary
Dataset automatically created during the evaluation run of model Intel/neural-chat-7b-v3-1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Intel__neural-chat-7b-v3-1.details_one-man-army__una-neural-chat-v3-3-P2-OMA
Dataset Card for Evaluation run of one-man-army/una-neural-chat-v3-3-P2-OMA
Dataset automatically created during the evaluation run of model one-man-army/una-neural-chat-v3-3-P2-OMA on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_one-man-army__una-neural-chat-v3-3-P2-OMA.details_Azure99__blossom-v5.1-34b
Dataset Card for Evaluation run of Azure99/blossom-v5.1-34b
Dataset automatically created during the evaluation run of model Azure99/blossom-v5.1-34b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Azure99__blossom-v5.1-34b.DetailVerifyBench
DetailVerifyBench
Project Page | Paper | GitHub
DetailVerifyBench is a rigorous benchmark designed for dense hallucination localization in long image captions. It comprises 1,000 high-quality images across five distinct domains: Chart, Movie, Nature, Poster, and UI. With an average caption length of over 200 words and dense, token-level annotations of multiple hallucination types, it stands as a challenging benchmark for evaluating the precise hallucination localization capabilities… See the full description on the dataset page: https://huggingface.co/datasets/zyxhhnkh/DetailVerifyBench.details_AA051610__FT
Dataset Card for Evaluation run of AA051610/FT
Dataset automatically created during the evaluation run of model AA051610/FT on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AA051610__FT.details_meta-llama__Llama-2-13b-hf
Dataset Card for Evaluation run of meta-llama/Llama-2-13b-hf
Dataset Summary
Dataset automatically created during the evaluation run of model meta-llama/Llama-2-13b-hf on the Open LLM Leaderboard.
The dataset is composed of 123 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 8 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_meta-llama__Llama-2-13b-hf.details_Qwen__Qwen2-1.5B-Instruct
Dataset Card for Evaluation run of Qwen/Qwen2-1.5B-Instruct
Dataset automatically created during the evaluation run of model Qwen/Qwen2-1.5B-Instruct.
The dataset is composed of 117 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/amztheory/details_Qwen__Qwen2-1.5B-Instruct.
