datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
za-african-next-voices
Swivuriso: ZA-African Next Voices
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech, collected through ethical, community-centered processes.
Dataset Paper: ArXiv - Work in Progress
Language Coverage… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.Nexora-music-pd-v1-mediumnexa-audiolm-instuct-benchmarkBEANS-Next
BEANS-Next
Audio files live under audio/. Metadata is in metadata.parquet with a
file_name column (repo-relative paths) so the Hugging Face Dataset Viewer can
play clips—only file_name uses that convention; tier 4 rows instead use
context_audio_paths (list) and query_audio_path so the viewer is not confused
by several *_file_name-like columns. Column task is the benchmark task id
(same strings as the old subset column in other layouts). Column tier is an
integer in {1,2,3,4}… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/BEANS-Next.214-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Dataset
Description
주제 기반 대화로 수집한 한국어(대한민국) 멀티스트림 자연 발화 스마트폰 음성 데이터셋입니다. 각 음성 데이터에는 대화 내용을 전사한 텍스트와 화자 ID, 성별, 연령 등의 속성 정보가 포함되어 있습니다. 본 데이터셋은 다양한 지역과 배경을 가진 폭넓은 화자들로부터 수집되었으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상에 활용할 수 있습니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1704?source=hf.kr
Specifications
Format
16 kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/214-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Dataset.za-african-next-voices-compressedNote: This dataset is a compressed version of za-african-next-voices. It was compressed to .opus format using a 32k bitrate.
Swivuriso: ZA-African Next Voices-Compressed
Swivuriso is a large-scale multilingual speech dataset targeting over 3000 hours of audio across 7 South African languages. The dataset is developed to support Automatic Speech Recognition (ASR) and inclusive speech technologies for low-resource African languages. It combines both scripted and unscripted speech… See the full description on the dataset page: https://huggingface.co/datasets/dsfsi-anv/za-african-next-voices-compressed.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample
Description
Arabic(Saudi) Multi-stream Spontaneous Dialogue Smartphone speech dataset-Customer Service. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1627?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample.500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset
Description
일본어(Japan) 48kHz 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 자유롭게 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다. 또한 다양한 AI 기업을 통해 데이터 품질을 검증했습니다.
데이터 수집, 저장 및 활용 전 과정에서 개인정보 보호 및 관련 법규를 엄격하게 준수하며, 사용자의 개인정보와 법적 권리를 보호합니다. 본 데이터셋은 GDPR, CCPA, PIPL을 준수합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1971?source=hf.kr
Specifications… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.172-hours-American-English-Full-Duplex-Multi-Channel-Speech-DatasetDescription
주제 기반 대화로 수집한 미국 영어 멀티스트림 자연 발화 스마트폰 음성 데이터셋입니다. 각 음성 데이터에는 대화 내용을 전사한 텍스트와 화자 ID, 성별, 연령 등의 속성 정보가 포함되어 있습니다. 본 데이터셋은 다양한 지역과 배경을 가진 폭넓은 화자들로부터 수집되었으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상에 활용할 수 있습니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1770?fromCategoryPage=1?source=hf.kr
Specifications
Format
16 kHz, 16 bit, 비압축 WAV, 모노 채널, 화자별 채널 분리
Content category
주제 기반 자연 대화
Recording condition
낮은 배경 소음 환경(실내)
Recording device
Android 스마트폰, iPhone
Country… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/172-hours-American-English-Full-Duplex-Multi-Channel-Speech-Dataset.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset
Description
사우디아라비아 아랍어(Arabic-Saudi) 멀티스트림 자연 대화 스마트폰 고객 서비스 음성 데이터셋입니다. 다양한 고객 서비스 상황에서 자유롭게 대화하는 방식으로 수집되었으며, 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 268명의 아랍어 원어민 화자로부터 데이터를 수집하여 다양한 화자 특성을 반영했으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1627?source=hf.kr
Specifications
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset.211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset
Description
211시간 규모의 태국어(Thai-Thailand) 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음했습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 654명의 태국어 원어민 화자로부터 데이터를 수집했으며, 다양한 화자와 지역적 특성을 반영하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1594?source=hf.kr
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.423-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Dataset
Description
필리핀 영어(English-Philippine) 멀티스트림 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1771?source=hf.kr
Specifications
Format
16kHz, 16 bit, 비압축 WAV, 모노 채널, 화자별 채널 분리
Content category
주어진 주제를 기반으로 대화
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/423-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Dataset.INTERSPEECH-2025-MLC-SLM-Challenge-Data-Sample
Description
The INTERSPEECH 2025 MLC-SLM Challenge Dataset, curated by Datatang, is derived from fifteen proprietary conversational speech corpora. Distinguished by exceptional annotation accuracy and operational reliability, this dataset is engineered to address critical challenges in multilingual automatic speech recognition (ASR) and long-context comprehension. It meticulously replicates real-world complexities including spontaneous interruptions and speaker overlaps across… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/INTERSPEECH-2025-MLC-SLM-Challenge-Data-Sample.NEX0211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample
Description
211 Hours - Thai(Thailand) Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(654 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.1100-Hours-Tagalog-Full-Duplex-Multi-Channel-Speech-Data-Sample
1100-Hours-Tagalog-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
Tagalog Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/1100-Hours-Tagalog-Full-Duplex-Multi-Channel-Speech-Data-Sample.200-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
Korean(Korea) Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1704?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/200-Hours-Korean-Full-Duplex-Multi-Channel-Speech-Data-Sample.600-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
English(Philippine) Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1771?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/600-Hours-English-Philippine-Full-Duplex-Multi-Channel-Speech-Data-Sample.600-Hours-American-English-Full-Duplex-Multi-Channel-Speech-Data-Sample
Description
American English Multi-stream Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks.
For
more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1770?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/600-Hours-American-English-Full-Duplex-Multi-Channel-Speech-Data-Sample.352-Hours-Urdu-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample
Description
Urdu Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/352-Hours-Urdu-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.nextmetal-lol
League of Legends Gameplay Dataset Sample
Description
This dataset sample contains high-quality screen captures of League of Legends gameplay, synchronized with corresponding keyboard and mouse input data. It serves as a specialized dataset for training long-horizon robotic agents. The core objective of this data is to provide a complex, multi-step interaction environment where agents can learn high-level planning and execution. These learned behaviors are… See the full description on the dataset page: https://huggingface.co/datasets/fiveoceans-dev/nextmetal-lol.beans-next-small
BEANS-Next
Audio files live under audio/. Metadata is in metadata.parquet with a
file_name column (repo-relative paths) so the Hugging Face Dataset Viewer can
play clips—only file_name uses that convention; tier 4 rows instead use
context_audio_paths (list) and query_audio_path so the viewer is not confused
by several *_file_name-like columns. Column task is the benchmark task id
(same strings as the old subset column in other layouts). Column tier is an
integer in {1,2,3,4} (legacy… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/beans-next-small.CATS-ami-next-speaker-audioNexora-music-pd-v1-mini
Nexora-Music-PD v1-mini
Overview
Nexora-Music-PD v1-mini is a curated public-domain music dataset composed of
historical audio recordings sourced from the Library of Congress Citizen DJ
collections. The dataset is designed for open research, audio analysis,
music information retrieval, remixing, and AI/ML experimentation.
This is a mini release (v1) intended as a lightweight, easy-to-use subset
for testing pipelines, educational use, and small-scale experiments.… See the full description on the dataset page: https://huggingface.co/datasets/ArkAiLab-Adl/Nexora-music-pd-v1-mini.nexa-audiolm-benchmarkNexus_image6
Bhagavad-Gita_Audio
Dataset Summary
Bhagavad-Gita_TTS is a high-quality, verse-aligned audio dataset of the Bhagavad Gita, designed for Text-to-Speech (TTS), Automatic Speech Recognition (ASR), Sanskrit NLP, and spiritual audio research.
Each shloka is paired with:
Original Sanskrit text
IAST transliteration
Clean High Quality 44.1KHz WAV audio recordings
The dataset is structured to mirror the 18 chapters of the Gita, covering all 701 shlokas… See the full description on the dataset page: https://huggingface.co/datasets/Sarkhan222/Nexus_image6.train_set_next_2000_en_de
Dataset Card for "train_set_next_2000_en_de"
More Information needed
tts_dataNExT-QA
