datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ioai-2026-find-the-order
IOAI 2026 — Find the Order
Official contest data for Find the Order, task 1 (Day 1) of the IOAI 2026 Individual Contest, held in Astana, Kazakhstan.
Spoken English dialogues are segmented into speaker turns, one .wav per turn, and shuffled. Contestants reconstruct the original turn order. prefix.json in each dialogue folder gives the first two chunk indexes, fixing the reading direction.
The task statement, translations, baseline and grader live in the IOAI-2026 GitHub… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/ioai-2026-find-the-order.toefl-2026-official-listening-audio
TOEFL iBT 2026 新格式 · 官方练习卷听力音频(镜像)
196 个音频文件 / 28 MB。取自 ETS 官方 7 套「对版练习卷」的听力部分,
文件名摊平以便直接按 URL 引用。
这是什么,不是什么
是:ETS 为 2026-01-21 起的新格式自己编写的官方练习材料。
每份 PDF 首页原文:
"This practice test aligns with TOEFL iBT tests from January 21, 2026.
It is not an exact replica of the actual test; directions and questions
have been adapted for paper format usability."
不是退役真题(不是 TPO 那种回收考卷),也不是任何一场真实考试的题目。
截至目前公开渠道没有 2026 新格式的退役真题,这是唯一 100% 对版的官方材料。
版权与来源
Copyright ©… See the full description on the dataset page: https://huggingface.co/datasets/xxfasdf/toefl-2026-official-listening-audio.urgent2024_official
Dataset Description
This dataset contains the official validation, non-blind test and blind test data used during the 2024 URGENT Speech Enhancement Challenge (https://urgent-challenge.github.io/urgent2024/), an official NeurIPS 2024 Competition. The dataset is designed to evaluate the performance of speech enhancement systems in handling various distortions and input conditions (e.g., varying sampling frequencies and encoding formats such as MP3/FLAC).
The dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/urgent-challenge/urgent2024_official.b150_official_test_100
B150 Public Eval Smoke 100
This dataset is a 100-sample public-eval smoke split released as raw
audio plus a lightweight JSONL manifest.
It is intended for low-cost pipeline bring-up and local fine-tuning
smoke tests. It is not the recommended split for formal experiments.
What Is Included
raw/audio/...: the released MP3 audio files
selected_samples.exported.jsonl: sample manifest with
audio_path, target_text, duration, language, and
prompt_boundary_index… See the full description on the dataset page: https://huggingface.co/datasets/maimai11/b150_official_test_100.viet-med-official-enviet-med-en-officialViet-med-official
