datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arena-resultsThis dataset contains the saved results from MTEB-Arena
mosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.ucf-crimeMultiTurn-Chat-MT-Bench-Judge
SEA-MT-Bench-Judge
SEA-MT-Bench-Judge expands on the original SEA-MTBench through the use of a criteria-based evaluation framework. We use GPT-OSS-120B as the judge model.
The prompts are based on MT-Bench and was manually translated by native speakers. Furthermore, some prompts were modified to be more suitable for the criteria-based judgments.
Supported Tasks and Leaderboards
SEA-MT-Bench-Judge is designed for evaluating chat or instruction-tuned large language… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/MultiTurn-Chat-MT-Bench-Judge.mt_bench_human_judgments
Content
This dataset contains 3.3K expert-level pairwise human preferences for model responses generated by 6 models in response to 80 MT-bench questions.
The 6 models are GPT-4, GPT-3.5, Claud-v1, Vicuna-13B, Alpaca-13B, and LLaMA-13B. The annotators are mostly graduate students with expertise in the topic areas of each of the questions. The details of data collection can be found in our paper.
Agreement Calculation
This Colab notebook shows how to compute the… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/mt_bench_human_judgments.mtnwx-trainingMagpie-Llama-3.1-Pro-MT-300K-Filtered
Project Web: https://magpie-align.github.io/
Arxiv Technical Report: https://arxiv.org/abs/2406.08464
Codes: https://github.com/magpie-align/magpie
Abstract
Click Here
High-quality instruction data is critical for aligning large language models (LLMs). Although some models, such as Llama-3-Instruct, have open weights, their alignment data remain private, which hinders the democratization of AI. High human labor costs and a limited, predefined scope for prompting prevent… See the full description on the dataset page: https://huggingface.co/datasets/Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered.tatdqa_test_beirBEIR version of vidore/tatdqa_test.
docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled.
infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled.
mteb-pt-results
🇧🇷 MTEB-BR — Benchmark Results
Canonical results store for MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark.
93 models · 22 native PT-BR tasks · 7 categories · no machine translation
What is this?
This repository is the canonical, machine-readable results store for MTEB-BR — a benchmark that evaluates text-embedding models on native Brazilian Portuguese (data created or found in Portuguese; machine-translated corpora such as… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/mteb-pt-results.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_healthcare_industry_test_beirBEIR version of vidore/syntheticDocQA_healthcare_industry_test.
arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled.
tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled.
syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
mTSBench
mTSBench
mTSBench is a collection of 344 multivariate time series from 19 datasets commonly used in anomaly detection research. Each folder corresponds to one dataset and contains *_train.csv, *_test.csv, and *_val.csv files. See data_summary.csv for per-file statistics.
How to download
This repository uses Git LFS for the CSV files.
git lfs install
git clone https://huggingface.co/datasets/PLAN-Lab/mTSBench
Load with Hugging Face
Select one of the… See the full description on the dataset page: https://huggingface.co/datasets/PLAN-Lab/mTSBench.syntheticDocQA_energy_test_beirBEIR version of vidore/syntheticDocQA_energy_test.
shiftproject_test_beirBEIR version of vidore/shiftproject_test.
MT_ColorThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/MT_Color.qwen36-kquant-offload-mtp-swebench-lite100-results
Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results
This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer.
Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control.
The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.cmath-mtThis dataset is a machine translated version of weitianwen/cmath.
Translated using dataset-translator.
dengue-mt-medallionMT-Reasoning
MultiSynt
MultiSynt is an open multilingual synthetic dataset.
The MT Reasoning subset of MultiSynt is made of automatic translations into 2 languages of Glaive AI reasoning dataset containing 22mil+ general reasoning questions, reasoning traces and responses.
lang
rows
prompt_tokens
reasoning_tokens
response_tokens
total_tokens
deu_Latn
17_354_716
1_873_153_732
26_010_932_738
14_862_651_336
42_746_737_806
fra_Latn
17_354_716
1_802_885_115
25_224_272_259… See the full description on the dataset page: https://huggingface.co/datasets/MultiSynt/MT-Reasoning.MT_Size_RecognitionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 46075,
"total_tasks": 6,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/MT_Size_Recognition.svq
Simple Voice Questions
Simple Voice Questions (SVQ) is a set of short audio questions recorded in 26 locales across 17 languages under multiple audio conditions.
Data Collection
Speakers were presented with recording instructions specifying the recording environment and text query to be recorded.
They recorded using their own phones or tablets under four conditions:
clean: Record in quiet environment
background speech noise: Record while audio from sources like podcasts… See the full description on the dataset page: https://huggingface.co/datasets/mteb/svq.AMPBench-MT
AMPBench-MT
AMPBench-MT is a homology-controlled benchmark for antimicrobial peptide endpoint prediction. The release is dated 2026-07-08.
Repository: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT
The benchmark is organized around endpoint-aware prediction rather than binary AMP recognition alone. It contains processed task tables for AMP/non-AMP classification, species-conditioned MIC regression, activity spectrum positive-evidence audits, low-toxicity classification… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.metaworld_mt10This dataset was created using LeRobot.
Dataset Description
NOTE:
All expert trajectories (100% success rate)
50 total episodes
Camera view: 3rd-person Corner2 only out of ["corner", "corner2", "corner3", "topview", "behindGripper"]
Generator script can be found here: https://github.com/aadarshram/lerobot/blob/MultiTask/src/lerobot/scripts/generate_MetaWorld_datasets.py
Homepage: [More Information Needed]
Paper: [More Information Needed]
License: apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/metaworld_mt10.MTID
TurnGate: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Overview
TurnGate is a response-aware defense mechanism designed to detect and mitigate hidden malicious intent in multi-turn dialogue systems. Defending state-of-the-art multi-turn malicious attacks like CKA-Agent.
MTID Dataset
We include the MTID (Multi-Turn Intent Dataset) here. This dataset contains a collection of… See the full description on the dataset page: https://huggingface.co/datasets/Graph-COM/MTID.MTP
