CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EchoSafe-MLLM /MM-SafetyBench-plus-plus MM-SafetyBench++ Project Page | Paper | Code MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent. Dataset Summary For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.imageimage-text-to-text1K<n<10K2 likes817 downloads6mo agoHugging Face02KurtDu /EchoX-Dialogues-Plus EchoX-Dialogues-Plus: Training Data Plus for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs 🐈‍⬛ Github | 📃 Paper | 🚀 Space  🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues (base)  EchoX-Dialogues-Plus EchoX-Dialogues-Plus extends KurtDu/EchoX-Dialogues with large-scale Speech-to-Speech (S2S) and Speech-to-Text (S2T) dialogues. All assistant/output speech is synthetic (single, consistent timbre for S2S). Texts are from… See the full description on the dataset page: https://huggingface.co/datasets/KurtDu/EchoX-Dialogues-Plus.automatic-speech-recognition1M<n<10M5 likes357 downloads1y agoHugging Face03FreedomIntelligence /EchoX-Dialougues EchoX-Dialogues: Training Data for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs 🐈‍⬛ Github | 📃 Paper | 🚀 Space  🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues-Plus  EchoX-Dialogues provides the primary speech dialogue data used to train EchoX, restricted to S2T (speech → text) in this repository. All input speech is synthetic; text is derived from public sources with multi-stage cleaning and rewriting. Most turns include asr /… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues.automatic-speech-recognition4 likes244 downloads1y agoHugging Face04Echo9k /mexican-legal-benchmarks Mexican Legal Benchmarks: Interpretation, Reasoning, and Cross-Jurisdiction Evaluation Dataset Summary The first specialized benchmark suite for evaluating language models on Mexican legal tasks. Three configs test distinct legal capabilities: practical interpretation of federal statutes, IRAC-structured reasoning chains with citation verification, and cross-jurisdiction comparison between Mexican states. 380 total samples across 3 benchmarks, drawn from Mexican federal… See the full description on the dataset page: https://huggingface.co/datasets/Echo9k/mexican-legal-benchmarks.texttext-classificationn<1K1 likes38 downloads5mo agoHugging Face05exnivo /Echo88-Instruct-173K Echo88 Instruct 173K A 173K-row retro instruction-tuning dataset for training Echo88-style small language models. Echo88 Instruct 173K is an English supervised fine-tuning dataset created for training exnivo/Echo88-150M-Instruct, the instruction-following version of Echo88. The dataset was built to teach a small language model how to answer questions, follow prompts, and behave like a helpful retro computer assistant whose knowledge is grounded in text from the 1950s… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/Echo88-Instruct-173K.texttext-generation100K<n<1M0 likes33 downloads3mo agoHugging Face06Vikaspandey582003 /echobench EchoBench The first public benchmark for LLM metacognitive calibration. EchoBench contains questions across 7 domains for training and evaluating whether language models accurately predict their own probability of being correct. Domains Domain Source Description Math GSM8K Grade-school math word problems Logic AI2-ARC Multiple-choice science reasoning Factual TriviaQA Open-domain factual questions Science SciQ Multiple-choice science questions Medical… See the full description on the dataset page: https://huggingface.co/datasets/Vikaspandey582003/echobench.textquestion-answering1K<n<10K0 likes23 downloads5mo agoHugging Face07ruohao /EchoMistgated Dataset Card for EchoMist Introducing EchoMist, the first comprehensive benchmark to measure how LLMs may inadvertently Echo and amplify Misinformation hidden within seemingly innocuous user queries. Dataset Description Prior work has studied language models' capability to detect explicitly false statements. However, in real-world scenarios, circulating misinformation can often be referenced implicitly within user queries. When language models tacitly agree, they may… See the full description on the dataset page: https://huggingface.co/datasets/ruohao/EchoMist.tabulartext-generationn<1K3 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.