adalat-ai/in22-legal
IN22-Legal Test-only out-of-distribution legal-domain dictation benchmark for Indic ASR. Read-speech recordings of legal passages from the IN22-Gen corpus, dense in domain entities (statute names, section numbers), formal numerals (dates, monetary amounts), and complex clause structures. Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (Interspeech 2026, under review). 📄 Documentation: DATASHEET.md… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/in22-legal.
Update datasheet
Link datasheet from README
Update source corpus reference
Add inline links
Add datasheet
Update transcripts
Update transcripts
Update transcripts
Update transcripts
Update transcripts
Per-language exact totals: audio duration, mean clip length, speaker count
Fix audio sample rate (48 kHz, not 16 kHz)
Add dataset card (SCRIBE release)
Upload dataset
Upload dataset
Upload dataset
initial commit
