CoolFace
Datasetpublic

adalat-ai/in22-legal

IN22-Legal Test-only out-of-distribution legal-domain dictation benchmark for Indic ASR. Read-speech recordings of legal passages from the IN22-Gen corpus, dense in domain entities (statute names, section numbers), formal numerals (dates, monetary amounts), and complex clause structures. Released alongside the SCRIBE evaluation framework in SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR (Interspeech 2026, under review). 📄 Documentation: DATASHEET.md… See the full description on the dataset page: https://huggingface.co/datasets/adalat-ai/in22-legal.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes42downloads
17 commits on main
9db98714mo ago

Update datasheet

kavyamanohar
2d863cf4mo ago

Link datasheet from README

kavyamanohar
7b47c3b4mo ago

Update source corpus reference

kavyamanohar
58417d94mo ago

Add inline links

kavyamanohar
e0de80a4mo ago

Add datasheet

kavyamanohar
ec13d234mo ago

Update transcripts

kavyamanohar
1d2cd174mo ago

Update transcripts

kavyamanohar
a565f975mo ago

Update transcripts

kavyamanohar
14b731e5mo ago

Update transcripts

kavyamanohar
620d4325mo ago

Update transcripts

kavyamanohar
31cb1a75mo ago

Per-language exact totals: audio duration, mean clip length, speaker count

kavyamanohar
024737b5mo ago

Fix audio sample rate (48 kHz, not 16 kHz)

kavyamanohar
109de585mo ago

Add dataset card (SCRIBE release)

kavyamanohar
6ccd0125mo ago

Upload dataset

kavyamanohar
3e57b855mo ago

Upload dataset

kavyamanohar
a4f6d0f5mo ago

Upload dataset

kavyamanohar
0eeb57c5mo ago

initial commit

kavyamanohar