liva-ai/hindi-english-asr
Hindi-English Code-Switching Conversational Audio This dataset contains conversational Hindi-English code-switching audio recordings with human-verified transcripts. The conversations feature natural, spontaneous speech between multiple speakers who fluidly switch between Hindi and English, a common pattern in urban South Asian speech communities. Transcription Process All transcripts go through at least two full human review passes. First, a native-speaker… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/hindi-english-asr.
Hindi-English Code-Switching Conversational Audio
This dataset contains conversational Hindi-English code-switching audio recordings with human-verified transcripts. The conversations feature natural, spontaneous speech between multiple speakers who fluidly switch between Hindi and English, a common pattern in urban South Asian speech communities.
Transcription Process
All transcripts go through at least two full human review passes. First, a native-speaker transcriber reviews and revises an AI-generated transcript using our custom-built transcript editor. Transcribers are provided the conversation as channel-separated audio, which enables precise speaker diarization correction and makes it easier to identify and label each speaker accurately. They work through the entire audio file to produce a full verbatim transcript with precise timestamps down to the millisecond, labeled speaker turns for diarization, and audio event tags for non-speech sounds such as [phone buzzing], [laughing], or [door closing]. Native fluency allows transcribers to accurately capture code-switching, overlaps, disfluencies, colloquialisms, and other conversational details that are difficult to capture without deep familiarity with the speakers and language variety.
Once the first pass is complete, the file is handed off to a senior reviewer with a proven track record of producing high-quality transcripts. The senior reviewer listens through the full audio file again and manually corrects any remaining transcription errors, spelling issues, timestamp inconsistencies, speaker label issues, or missed audio events. In addition, a senior native project lead performs rigorous spot checks across completed files to monitor quality, enforce consistency, and identify any recurring issues that need to be corrected across the project.
Timecodes
Transcripts are labeled as either overlapping or non-overlapping based on their timecode structure. In overlapping transcripts, speaker turns may have timestamps that overlap, capturing moments where multiple speakers talk simultaneously. In non-overlapping transcripts, timecodes are strictly sequential — no two speaker turns share the same time range — though the underlying audio may still contain overlapping speech. We provide transcripts tailored to any style guide, and our customers have requested a wide range of formatting and annotation conventions.
