datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.field-notes-transcription-packet
Field-notes Transcription Packet
Verified field-note transcriptions retained for the research archive.
Retained notes: 6
Featured note: NOTE-1002 — Cedar / English
Transcription window: 2023-01-18T16:20:00Z to 2023-12-15T10:00:00Z
Site counts (Cedar/Lark/Morrow/Pine): 2/2/1/1
Mean word counts (Cedar/Lark/Morrow/Pine): 422.5/362.5/360.0/480.0
ipa-transcription-datase
🗣️ English Text → IPA Transcription Dataset
Overview
This dataset provides a large-scale, phonemically rich collection of English text paired with International Phonetic Alphabet (IPA) transcriptions, designed to support research and applications in speech-language pathology, phonetics, and natural language processing.
It was created to enable data-driven phonetic transcription, reducing reliance on traditional rule-based systems and supporting modern… See the full description on the dataset page: https://huggingface.co/datasets/dsvv-cair/ipa-transcription-datase.luciolescribe-transcription-faq
🎙️ LucioleScribe Transcription FAQ - Dataset Production v2.0
📋 Description
Dataset enrichi de questions-réponses FAQ sur la transcription IA 100% locale avec LucioleScribe, première plateforme française de transcription conforme RGPD par conception.
🎯 Caractéristiques clés
📊 Taille: 182+ paires question-réponse (expansion continue)
🏷️ Métadonnées: Enrichi avec catégories, intents, buyer stages, styles
🌍 Langue: Français (France)
📑 Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/luciolescribe-transcription-faq.yt-transcriptionscantonese-youtube-transcription-fusionhttps://huggingface.co/datasets/alvanlii/cantonese-youtube 数据集中train-00000-of-01090.parquet 到 train-00350-of-01090.parquet 部分的转写文本清洗。使用qwen3-asr、qwen3-omni、sensevoicesmall(https://huggingface.co/ASLP-lab/WSYue-ASR)
进行转写,然后用 Qwen3.6-35B-A3B 根据语义进行转写纠正。
yt_transcriptionsqrl-yt-transcriptionsyoutube-transcriptionsendocrinology_transcription_and_notestranscriptions_formattedyoutube-transcriptionsyoutube-transcriptions-demotranscriptions
