datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HearInContextEnglish | 中文
HearInContext
A Benchmark for Implicit Context in Speech Recognition
Illustrative example: the same spoken request is disambiguated as flour or flower by different assistant histories. The dialogue and waveform are illustrative.
Same audio. Different contexts. Different meanings.
HearInContext is a Mandarin–English contextual speech recognition benchmark. It pairs the same audio with dialogue histories supporting different meanings to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/HearInContext.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.oral-arguments-us
US Court Oral Arguments -- metadata and transcripts
Code & tools: github.com/docketx — legal-scrambler pseudonymises a case file on your own hardware before a frontier model sees it; claude-for-legal is the Claude Code plugin (docketx-open-law) that loads these datasets and checks citations against them.
Source & credit — Free Law Project / CourtListener
Every recording catalogued here was collected and catalogued by CourtListener, and this dataset is
sliced… See the full description on the dataset page: https://huggingface.co/datasets/docketx/oral-arguments-us.lunde_nor_nob_reading_optimisedTest only - not for training.
First version - 0.1 of lunde_nor_nob_reading_optimised
This dataset does not contain any audio data.
Export Details
Train samples: 10040932
Validation samples: 0
Test samples: 0
Dataset created using search datasets:lunde_nor_nob_reading_optimised.
open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
12,000 chunks
30 seconds per chunk
100.0 hours total
Each item contains S1/S2 speaker labels, turn timings, target text,
relationship, pronouns, environment, topic, mood, and source reference IDs.
Audio renderer: vLLM-Omni VoxCPM2
Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.AI_Agent_Task_Dataset
🤖 Massive AI Agent Task Dataset (10.5GB)
📌 Overview
Welcome to the AI Agent Task Dataset, a massive 10.5GB procedural dataset designed for training, fine-tuning, and evaluating autonomous AI agents and LLMs.
This dataset focuses on:
Multi-step reasoning
Tool usage (APIs, frameworks, systems)
Real-world execution workflows
Perfect for building agentic AI systems, copilots, and automation models.
📑 Table of Contents
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/oncody/AI_Agent_Task_Dataset.asr-benchmark-outputs
SaarAI ASR Benchmark Outputs
Raw per-utterance model outputs (transcription manifests) produced by the gsma-asr-bench runners on SaarAI/asr-leaderboard-datasets.
files: 508
utterances: 4390208
languages: 7
models: 47
Layout
data/<language_name>/<split>__<dataset_config>__<model_slug>.jsonl
index.jsonl # one record per file (language, split, model, rows, sha256, ...)
index.csv
Directories categorise by language name; the file name begins with the split name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/asr-benchmark-outputs.mosla
Overview
The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one
of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction,
and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human
annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.nepal-oral-demo
Nepal Oral Demo
Public, always-safe fixtures for the Nepal oral-language backbone (AkAiNp).
Fictional “Demo Himalayan” track only
Schema examples for CI, export-script tests, and the Expo training app offline demo pack
No real community speakers, ever
Monorepo: nepal-multilingual-llm (local project). Source: data/packs/_demo/ + packages/schema/examples/.
Intended uses
OK
Not OK
Unit tests, pipeline dry-runs
Training production ASR/TTS as if it were… See the full description on the dataset page: https://huggingface.co/datasets/AkAiNp/nepal-oral-demo.
