CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IMJONEZZ /star-wars-dataset Star Wars dialogue dataset - annotation layers (snapshot 2026-09-21) One master cue table serving diarization, ASR and dialogue-LLM training: every subtitle cue part of 15 titles with a forced-aligned time span, a character label and provenance. No audio or video is included - the source media is copyrighted. Clip paths in the manifests resolve once you rebuild the clips from your own copies with the pipeline code (export_asr.py, export_diarization.py). The text layers contain… See the full description on the dataset page: https://huggingface.co/datasets/IMJONEZZ/star-wars-dataset.tabularautomatic-speech-recognition10K<n<100K3 likes83 downloads3d agoHugging Face02issai /KazMix-3 KazMix-3 Kazakh three-speaker overlapping-speech dataset for target-speaker ASR (TS-ASR), released with the Persona-ASR project. Given a short enrollment utterance of a target speaker and a 3-speaker mixture, the task is to transcribe only the target speaker, or reject the utterance when the target is absent. This repository ships the mixture manifests and generation scripts, not the audio. Mixtures are derived from the Kazakh Speech Dataset (KSD, OpenSLR 140); download KSD and… See the full description on the dataset page: https://huggingface.co/datasets/issai/KazMix-3.tabularautomatic-speech-recognition10K<n<100K0 likes41 downloads10d agoHugging Face03bdx33 /ted-talks-in-chinese-zhongwen TED中文 Podcast 聚焦华语地区的创意,本节目从上万个TED和TEDx演讲中,为您精选中文演讲,以及少量中文配音的经典英语演讲。演讲人包括科技和人文专家、关心当下与未来的思考者、关注挑战与探索的实践者。英雄不论出处,谁有创意谁讲。让这些演讲成为一把把钥匙,开启你的好奇心,升级你的行动力。 Focusing on creativity within the Chinese-speaking world, this program curates Chinese-language talks from tens of thousands of TED and TEDx presentations, along with a select few English classics dubbed into Chinese. Our speakers span technology and humanities experts, thinkers engaged with the present and future, and practitioners… See the full description on the dataset page: https://huggingface.co/datasets/bdx33/ted-talks-in-chinese-zhongwen.texttext-generationn<1K1 likes28 downloads1y agoHugging Face04IDEMITSU /hoiku-yougo-stt-ja 保育用語STT最適化データセット Japanese Childcare Vocabulary Dataset for STT Optimization 🇯🇵 日本語 概要 このデータセットは、日本の保育現場における音声認識(STT)の精度向上を目的として構築された、保育専用の構造化語彙集です。 一般的なSTTエンジン(Google・Amazon・OpenAI Whisper等)は、保育現場で日常的に使われる専門用語を正確に認識できないという問題があります。たとえば: 実際の発話 STTが誤認識する例 誤飲(ごいん) 語韻・誤音 午睡(ごすい) 誤推・誤睡 降園(こうえん) 公演・後援 点呼(てんこ) 転校・天候 視診(ししん) 試診・指診 このデータセットは、こうした保育特有の同音異義語・誤変換パターンを体系的に収録し、STTエンジンのカスタム辞書・RAG(検索拡張生成)・AIコール(音声電話応答)に活用できるよう設計されています。… See the full description on the dataset page: https://huggingface.co/datasets/IDEMITSU/hoiku-yougo-stt-ja.textautomatic-speech-recognitionn<1K0 likes16 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.