datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OCRBenchGithub|Paper
OCRBench has been accepted by Science China Information Sciences.
EchoPolicy-0-MolmoSpaces-eval
EchoPolicy-0: MolmoSpaces Evaluation
Evaluation results, complete episode videos, trajectories and execution logs for EchoPolicy-0, developed by MagicLab.
Official submission: allenai/molmospaces#199
Evaluation version: echopolicy-ms-v15-20260918-full3915
Policy version: echo-tools-20260917-v15
Coverage: 3915 episodes across Close, Pick, Open and Pick & Place, including every success and failure.
Browse the videos
Open the episode viewer.
Each row represents one… See the full description on the dataset page: https://huggingface.co/datasets/ddffwyb/EchoPolicy-0-MolmoSpaces-eval.MM-SafetyBench-plus-plus
MM-SafetyBench++
Project Page | Paper | Code
MM-SafetyBench++ is a benchmark designed for evaluating contextual safety in Multi-Modal Large Language Models (MLLMs). It challenges models to distinguish subtle contextual differences between scenarios that may appear visually or textually similar but diverge significantly in safety intent.
Dataset Summary
For each unsafe image-text pair, the benchmark includes a corresponding safe counterpart created through minimal… See the full description on the dataset page: https://huggingface.co/datasets/EchoSafe-MLLM/MM-SafetyBench-plus-plus.llm2vec-gen-echo-rewritten-w-hard-negative
LLM2Vec-Gen
The dataset consists of generations based on the Echo data (Springer et al). The instruction+queries are rewritten in a natural tone using Gemini. The generations are intended to be used for training LLM2Vec-Gen models, serving as the target output for queries.
The negative_question in this dataset are also generated by Gemini. This dataset consists of various splits. Each split corresponds to responses generated by a specific LLM, e.g., Qwen3-4B. The "original" split… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/llm2vec-gen-echo-rewritten-w-hard-negative.EchoFake
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
Paper link: http://arxiv.org/abs/2510.19414
Code for baseline models is available at https://github.com/EchoFake/EchoFake
Auto-recording tools is available at https://github.com/EchoFake/EchoFake/tree/main/tools
Abstract
The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing… See the full description on the dataset page: https://huggingface.co/datasets/EchoFake/EchoFake.Echo-4o-Image-Surrel-FantasyEchoes-Platos-CaveEchoes in Plato's Cave — Controlled Speech–Text Corpus
Controlled corpus of 14,400 synthetic English utterances in which the same 600 sentences are
rendered by 6 speakers × 4 emotions, so that speaker identity and prosody vary
while linguistic content is held fixed. It was built for the paper: Echoes in Plato's Cave: Measuring Global and Local Alignment Between Speech and Language Representations, accepted as an oral presentation at the Speech and Audio Language… See the full description on the dataset page: https://huggingface.co/datasets/alefiury/Echoes-Platos-Cave.echo-clones-4m-en
echo-clones-4m-en
~4 M English TTS clone utterances generated with
EchoTTS (jordand/echo-tts-base).
Sample rate: 44 100 Hz, 16-bit PCM WAV stored in Parquet
Speakers: 4 000 reference speakers (spk_0000-spk_3999)
Text bucketing: quip (<=100 chars), mid (100-300), ramble (300-420)
Speaker assignment: round-robin -- text[i] -> spk_{i % 4000}
Companion datasets
Reference speakers: SynData-2/echo-ref-speakers-4k-en -- the 4 000 reference WAVs used as speaker… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-EN/echo-clones-4m-en.EchoEval
EchoEval
EchoEval is a instance-level spoken empathetic evaluation benchmark. It comprising 1K authentic recordings from 20 professional actors
Load one subset:
from datasets import load_dataset
ds = load_dataset("ddwang2000/EchoEval", "normal", split="test")
Subsets
Subset
Size
Description
normal
220
Explicit, everyday emotional delivery
implicit
220
Emotion is present but understated in the text
very_high_intense
220
High-arousal, strongly… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/EchoEval.EchoLens
EchoLens
A Human-Speech Dataset for Auditing Demographic Sensitivity in Audio-Language Models
📄 Paper (EMNLP 2026 Findings) ·
💻 Code
Voice interfaces are increasingly moving away from transcription pipelines toward end-to-end systems that directly respond to audio inputs. This development in turn requires a shift in evaluation methodology away from transcription accuracy and towards more substantive markers such as response validity. We introduce EchoLens, a demographically… See the full description on the dataset page: https://huggingface.co/datasets/alexsdl/EchoLens.Orbdivination
Thyme Video RL — pass-rate filtered subset
RL training data for a video agent that first skims 64 uniformly sampled frames,
then calls a temporal-crop tool to pull higher-FPS segments from the source mp4.
Each row therefore carries both:
videos — the 64 overview frames, embedded as JPEG bytes
video_path — the source mp4, which the crop tool reads at rollout time
Provenance
Built from 12,000 multiple-choice questions over 1,479 videos (LVBench,
LongVideoBench… See the full description on the dataset page: https://huggingface.co/datasets/EchoMinkki/Orbdivination.mbti-cleaned
Dataset Card for "mbti-cleaned"
More Information needed
HLE_mathEchoFake
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
Paper link: http://arxiv.org/abs/2510.19414
Code for baseline models is available at https://github.com/EchoFake/EchoFake
Auto-recording tools is available at https://github.com/EchoFake/EchoFake/tree/main/tools
Abstract
The growing prevalence of speech deepfakes has raised serious concerns, particularly in real-world scenarios such as telephone fraud and identity theft. While many anti-spoofing… See the full description on the dataset page: https://huggingface.co/datasets/nccm2p2/EchoFake.pause63-echo.notag.r-1.0-k-8.statml-arxivskill_profession_job_echo_just_one
Dataset Card for "skill_profession_job_echo_just_one"
More Information needed
echo-data-rewritten-queries-hard-negative-qwen3-4becho-data-rewritten-queriesecho-tts-en-benchmarks-v1
Echo-TTS English Benchmarks v1
Описание
Датасет содержит результаты бенчмарка модели Echo-TTS на 720 предложениях из Harvard Sentences.
Каждое предложение озвучено 14 спикерами — 5 встроенных голосов + 9 голосов клонированных из KaniTTS-2.
Структура
Поле
Тип
Описание
text
string
Текст предложения
echo_tts_af_bella
audio
Голос AF Bella (44100 Hz)
echo_tts_af_heart
audio
Голос AF Heart (44100 Hz)
echo_tts_am_fenrir
audio
Голос AM Fenrir (44100… See the full description on the dataset page: https://huggingface.co/datasets/data-lab-voice/echo-tts-en-benchmarks-v1.echo-data-rewritten-queries-hard-negative-qwen3-8bZambezi_ECHO_v1
Zambezi ECHO v1: Shona-English Code-Switched Maternal Health Queries
Dataset Description
Zambezi ECHO v1 SESB (Shona-English Speech Benchmark) is a dataset of short,
simulated patient voice queries in Shona (Zimbabwe), code-switched with
English, covering common maternal and child health concerns — pregnancy
symptoms, danger signs, child illness, and general health questions asked
the way patients actually phrase them in the field, mixing Shona with
English… See the full description on the dataset page: https://huggingface.co/datasets/dawahealth/Zambezi_ECHO_v1.Echoverse
Echoverse
Synthetic computer-use / web-agent evaluation environments. Each environment is a
self-contained world that drills specific UI skills; an agent is given a natural-language goal
and must drive the environment's UI to the target state, graded against a reference_state_change
or reference_answer.
This dataset bundles, for every environment:
<env>/test_tasks.parquet — the shipped test tasks (previewable in the dataset viewer, with
per-column statistics and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Echoverse.PanelTS-ICLR-Review
PanelTS full data release
This release includes all 1,844 CSV files from the public PanelTS source repository at revision bfb751a6735771b5da4e51de2d08fb49a1ad0036, plus 17 canonical selections / 19 Hugging Face configurations. Raw CSV bytes are preserved under PanelTS/; no source data were regenerated. Submission author names are omitted.
Completeness is relative to this pinned upstream snapshot. File counts are not counts of independent scientific datasets. Configurations may… See the full description on the dataset page: https://huggingface.co/datasets/Echo0623/PanelTS-ICLR-Review.skill_profession_job_echo
Dataset Card for "skill_profession_job_echo"
More Information needed
echo2025
ECHO Benchmark
This repository contains the dataset accompanying the paper Constantly Improving Image Models Need Constantly Improving Benchmarks.
Project page: https://echo-bench.github.io/
Code: https://github.com/para-lost/ECHO
For any questions or inquiries, please contact us at echo-bench@googlegroups.com.
About the Dataset
ECHO stands for Extracting Community Hatched Observations. ECHO is a framework for constructing benchmarks directly from social media… See the full description on the dataset page: https://huggingface.co/datasets/echo-bench/echo2025.Zambezi_ECHO_v1
Zambezi ECHO v1: Shona-English Code-Switched Maternal Health Queries
Dataset Description
Zambezi ECHO v1 SESB (Shona-English Speech Benchmark) is a dataset of short,
simulated patient voice queries in Shona (Zimbabwe), code-switched with
English, covering common maternal and child health concerns — pregnancy
symptoms, danger signs, child illness, and general health questions asked
the way patients actually phrase them in the field, mixing Shona with
English… See the full description on the dataset page: https://huggingface.co/datasets/tarirozw/Zambezi_ECHO_v1.echocho_20260801_174804This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/shuren2/echocho_20260801_174804.echo-review-vectors
echo-review-vectors
Precomputed embeddings for the 45,864 distinct review texts in
Echo, encoded with
aynaval2003/echo-sbert-domain.
file
what it is
sbert-domain.fp16.npy
(45864, 768) float16, L2-normalised, row i matches row i of the parquet
corpus.parquet
row_index, content, n_rows, review_ids
45,864 vectors cover 64,280 review rows, because identical texts share one vector.
corpus.parquet is what joins a vector back to its reviews — without it the .npy
is an… See the full description on the dataset page: https://huggingface.co/datasets/aynaval2003/echo-review-vectors.ECHO-Terminal-Agent-Prepared-Data
ECHO-style Terminal Agent Prepared Data for LFM RLVR
Prepared on 2026-06-09 for local no-Docker LFM terminal RLVR experiments.
This dataset converts public terminal-agent task archives into two formats:
echo_terminal_tasks_*.parquet: ECHO/SkyRL-style rows with prompt, path, and task_binary.
lfm_live_tasks_mixed.*: local LFM no-Docker trainer rows with prompt, task_id, source, task_binary_b64, and metadata.
Current manifest:
total rows: 1500
Endless Terminals: 772 rows… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/ECHO-Terminal-Agent-Prepared-Data.eng_echo
