datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.EchoX-Dialogues-Plus
EchoX-Dialogues-Plus: Training Data Plus for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues (base)
EchoX-Dialogues-Plus
EchoX-Dialogues-Plus extends KurtDu/EchoX-Dialogues with large-scale Speech-to-Speech (S2S) and Speech-to-Text (S2T) dialogues.
All assistant/output speech is synthetic (single, consistent timbre for S2S). Texts are from… See the full description on the dataset page: https://huggingface.co/datasets/KurtDu/EchoX-Dialogues-Plus.EchoX-Dialougues
EchoX-Dialogues: Training Data for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues-Plus
EchoX-Dialogues provides the primary speech dialogue data used to train EchoX, restricted to S2T (speech → text) in this repository.
All input speech is synthetic; text is derived from public sources with multi-stage cleaning and rewriting. Most turns include asr /… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues.mexican-legal-benchmarks
Mexican Legal Benchmarks: Interpretation, Reasoning, and Cross-Jurisdiction Evaluation
Dataset Summary
The first specialized benchmark suite for evaluating language models on Mexican legal tasks. Three configs test distinct legal capabilities: practical interpretation of federal statutes, IRAC-structured reasoning chains with citation verification, and cross-jurisdiction comparison between Mexican states.
380 total samples across 3 benchmarks, drawn from Mexican federal… See the full description on the dataset page: https://huggingface.co/datasets/Echo9k/mexican-legal-benchmarks.Echo88-Instruct-173K
Echo88 Instruct 173K
A 173K-row retro instruction-tuning dataset for training Echo88-style small language models.
Echo88 Instruct 173K is an English supervised fine-tuning dataset created for training exnivo/Echo88-150M-Instruct, the instruction-following version of Echo88.
The dataset was built to teach a small language model how to answer questions, follow prompts, and behave like a helpful retro computer assistant whose knowledge is grounded in text from the 1950s… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/Echo88-Instruct-173K.echobench
EchoBench
The first public benchmark for LLM metacognitive calibration.
EchoBench contains questions across 7 domains for training and evaluating
whether language models accurately predict their own probability of being correct.
Domains
Domain
Source
Description
Math
GSM8K
Grade-school math word problems
Logic
AI2-ARC
Multiple-choice science reasoning
Factual
TriviaQA
Open-domain factual questions
Science
SciQ
Multiple-choice science questions
Medical… See the full description on the dataset page: https://huggingface.co/datasets/Vikaspandey582003/echobench.EchoMist
Dataset Card for EchoMist
Introducing EchoMist, the first comprehensive benchmark to measure how LLMs may inadvertently Echo and amplify Misinformation hidden within seemingly innocuous user queries.
Dataset Description
Prior work has studied language models' capability to detect explicitly false statements. However, in real-world scenarios, circulating misinformation can often be referenced implicitly within user queries. When language models tacitly agree, they may… See the full description on the dataset page: https://huggingface.co/datasets/ruohao/EchoMist.
