datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.cite-llm-multi-cite-trainllmpicto-commonvoice-v2-multiAutomatic translation of benoitfavre/llmpicto-commonvoice-v2 from French to languages where the Arasaac lexicon is available. Generated with facebook/nllb-200-distilled-600m. Contains about 500k sentences for 35 languages.
details_MTSAIR__multi_verse_model
Dataset Card for Evaluation run of MTSAIR/multi_verse_model
Dataset automatically created during the evaluation run of model MTSAIR/multi_verse_model on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_MTSAIR__multi_verse_model.NASDAQ-News-Multi-LLM-Scores
NASDAQ News Multi-LLM Scores
127,176 financial news articles scored by 11 state-of-the-art LLMs for sentiment and risk assessment.
This dataset takes the same articles from FNSPID / FinRL_DeepSeek and re-scores them using multiple LLMs with varying reasoning effort levels and summary inputs. It enables direct cross-model comparison of financial sentiment analysis on identical articles.
Motivation
When we began using the FNSPID dataset for RL trading agent research, we… See the full description on the dataset page: https://huggingface.co/datasets/HYL/NASDAQ-News-Multi-LLM-Scores.details_ammarali32__multi_verse_model
Dataset Card for Evaluation run of ammarali32/multi_verse_model
Dataset automatically created during the evaluation run of model ammarali32/multi_verse_model on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_ammarali32__multi_verse_model.Kazakh_Multi-Task_corpus
Kazakh Multi-task Corpus
A multi-task NLP dataset in the Kazakh language, covering seven distinct language tasks - from instruction-following and question answering to translation, sentiment analysis, and grammar exercises. Designed to support the development of Kazakh-language models, benchmarks, and linguistic research.
Dataset Summary
Kazakh is a Turkic language spoken by over 13 million people, yet it remains significantly underrepresented in NLP research and… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/Kazakh_Multi-Task_corpus.details_Salesforce__codegen-6B-multi
Dataset Card for Evaluation run of Salesforce/codegen-6B-multi
Dataset Summary
Dataset automatically created during the evaluation run of model Salesforce/codegen-6B-multi on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Salesforce__codegen-6B-multi.details_Josephgflowers__GPT2-774M-CINDER-SHOW-MULTI-CHAT
Dataset Card for Evaluation run of Josephgflowers/GPT2-774M-CINDER-SHOW-MULTI-CHAT
Dataset automatically created during the evaluation run of model Josephgflowers/GPT2-774M-CINDER-SHOW-MULTI-CHAT on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Josephgflowers__GPT2-774M-CINDER-SHOW-MULTI-CHAT.details_saishf__Multi-Verse-RP-7B
Dataset Card for Evaluation run of saishf/Multi-Verse-RP-7B
Dataset automatically created during the evaluation run of model saishf/Multi-Verse-RP-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_saishf__Multi-Verse-RP-7B.IFAAB-MULTI-LLM-2026
Dataset Card for IFAAB-MULTI-LLM-2026
This dataset card serves as a comprehensive datasheet for the kmhj1306/IFAAB-MULTI-LLM-2026 dataset repository. It maps demographic persona features to localized automated financial planning prompts and responses, specifically curated to evaluate LLM behavior within the Indian socio-economic context.
Dataset Details
Dataset Description
This dataset consists of 222,138 rows of tabular text data designed to… See the full description on the dataset page: https://huggingface.co/datasets/kmhj1306/IFAAB-MULTI-LLM-2026.details_MaziyarPanahi__YamshadowInex12_Multi_verse_modelExperiment28
Dataset Card for Evaluation run of MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28
Dataset automatically created during the evaluation run of model MaziyarPanahi/YamshadowInex12_Multi_verse_modelExperiment28 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_MaziyarPanahi__YamshadowInex12_Multi_verse_modelExperiment28.harmful-prompts-multillm-before-guardrail-evaluationharmful-prompts-multillm-after-guardrail-evaluation-newmulti_llm_dpogn-multi-affective-alpaca
Guaraní Multidimensional Affective Alpaca Dataset
This dataset is a multidimensional affective instruction-tuning dataset for Guaraní (gn), created by preprocessing and unifying three original Guaraní classification datasets into the Alpaca format (instruction/input/output). It is designed for fine-tuning large language models (LLMs) on low-resource indigenous language tasks.
📚 Source Datasets
This dataset combines and transforms the following datasets from Marvin M.… See the full description on the dataset page: https://huggingface.co/datasets/Capibara-LLM/gn-multi-affective-alpaca.rebel_multi_llm_as_userharmful-prompts-multillm-after-guardrail-evaluationseed_math_multiple_samples_scale_up_scaredy_cat_multiple_samples_llm_verifier_multi_domaincite-llm-multi-cite-evalfft-multi-tool-router-datamultillm-route-instruct2harmful-prompts-multillm-after-guardrail-evaluation-newl1-10-multi-label-tagger-datamultillm-route-instructharmful-prompts-multillm-after-guardrail-evaluation-newwiki_rephrase_multillm_brazil
wiki_rephrase_multillm_brazil
Dataset em portugues derivado de conteudos da Wikipedia em portugues, com um subset
para os textos originais e subsets separados para reescritas sinteticas. A reescrita
foi feita utilizando quatro modelos diferentes de três famílias de llms:
google/gemma-4-E4B-it
google/gemma-4-E2B-it
mistralai/Ministral-3-3B-Instruct-2512
Qwen/Qwen3.5-4B
Repositorio alvo: wiki_rephrase_multillm_brazil.
Visibilidade pretendida no Hugging Face Hub: private.… See the full description on the dataset page: https://huggingface.co/datasets/br-llm-data/wiki_rephrase_multillm_brazil.virtuoussy_multi_subject_rlvr_llm_judgexlam_multi_turn_FUNreason_the_original_3epoch_gradient_accumulate32
