multi-llm
DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
Multi-turn_Long-context_Benchmark_for_LLMs
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Arxiv: https://www.arxiv.org/abs/2507.13681
Huggingface: https://huggingface.co/papers/2507.13681
Introduction
LoopServe Multi-Turn Dialogue Benchmark is a comprehensive evaluation dataset comprising multiple diverse datasets designed to assess large language model performance in realistic conversational scenarios.
Unlike traditional benchmarks that place queries only at the end… See the full description on the dataset page: https://huggingface.co/datasets/TreeAILab/Multi-turn_Long-context_Benchmark_for_LLMs.cite-llm-multi-cite-trainllmpicto-commonvoice-v2-multiAutomatic translation of benoitfavre/llmpicto-commonvoice-v2 from French to languages where the Arasaac lexicon is available. Generated with facebook/nllb-200-distilled-600m. Contains about 500k sentences for 35 languages.
details_MTSAIR__multi_verse_model
Dataset Card for Evaluation run of MTSAIR/multi_verse_model
Dataset automatically created during the evaluation run of model MTSAIR/multi_verse_model on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_MTSAIR__multi_verse_model.NASDAQ-News-Multi-LLM-Scores
NASDAQ News Multi-LLM Scores
127,176 financial news articles scored by 11 state-of-the-art LLMs for sentiment and risk assessment.
This dataset takes the same articles from FNSPID / FinRL_DeepSeek and re-scores them using multiple LLMs with varying reasoning effort levels and summary inputs. It enables direct cross-model comparison of financial sentiment analysis on identical articles.
Motivation
When we began using the FNSPID dataset for RL trading agent research, we… See the full description on the dataset page: https://huggingface.co/datasets/HYL/NASDAQ-News-Multi-LLM-Scores.
