PrathamOrgAI/ReadNet
ReadNet Dataset Description ReadNet is an audio dataset collected from more than 2,00,000 children in the age group of 5-16 years in Hindi and Marathi language. The dataset consists of audio files in the wav format where children read out ASER Samples which consists of letters, words, stories and paragraphs in their native language. This dataset is a subset of the larger dataset which consists of an estimated 2500 hours of data. This dataset consists of ~87 hours… See the full description on the dataset page: https://huggingface.co/datasets/PrathamOrgAI/ReadNet.
0253
1---2license: cc-by-nc-4.03task_categories:4- automatic-speech-recognition5language:6- hi7- mr8---9## Dataset Description10 11- **Homepage:** [Pratham Education Foundation](https://www.pratham.org/)12- **Point of Contact:** [Satish Kumar](mailto:satish.k@pratham.org)13# ReadNet14## Dataset Description15ReadNet is an audio dataset collected from more than 2,00,000 children in the age group of 5-16 years in Hindi and Marathi language.16The dataset consists of audio files in the wav format where children read out [ASER Samples](https://asercentre.org/wp-content/uploads/2022/12/Hindi_ASER-2018.pdf) which consists of letters, words, stories and paragraphs in their native language.17This dataset is a subset of the larger dataset which consists of an estimated 2500 hours of data. This dataset consists of ~87 hours of audio data in Hindi and Marathi,18### Languages19Hindi and Marathi20```python21from datasets import load_dataset22 23dataset = load_dataset("PrathamOrgAI/ReadNet")24```25 26### Data Fields27 28- URL: URL of the audio file which can be downloaded in the .wav format29 30- Transcribed Text: Transcription of audio files in their respective language31 32 33### Annotations34Annotations for the audio files are provided in the Transcribed Text Column. The annotation of audio files was done on an annotation portal which was developed inhouse by Pratham Education Foundation35 36#### Annotation process37 38A team of annotators were hired for the annotation of audio39samples in two languages. Training regarding the annotation40portal was provided to the annotators. A dry run for the an-41notation was done initially and the annotation was reviewed42and feedback was given to the annotators to remove incon-43sistencies and errors in annotation. Annotation guidelines are44provided in the annotation portal itself, in both languages for45the annotators reference. Since the scale of the data set is46pretty large, we have decided to annotate some portion of the47data twice and the remaining data once depending on the re-48sources49 50#### Who are the annotators?51 52The Annotation was done by experts from the ASER team 53 54### Personal and Sensitive Information55 56The dataset consists of people who have donated their voice online. You agree to not attempt to determine the identity of speakers in this dataset.57 58 59### Social Impact of Dataset60The ReadNet Dataset is one of the largest corpus of children's speech in Hindi and Marathi Language. The Dataset was created with an aim to develop custom speech recognition models for assessing reading levels of children. The dataset can help in the development of automated tools for assessment of reading ability among children in Hindi and Marathi61### Licensing Information62 63Public Domain, Creative Commons Attribution 4.0 International Public License ([CC-BY-NC-4.0](https://creativecommons.org/licenses/by-nc/4.0/legalcode.en))64### Acknowledgements65 66We would like to thank Schmidt Futures and the Sarva Mangal Family trust for funding the ReadNet Project. We would67also like to thank Dr. Wilma Wadhwa and Anil Kumar Kamath from the ASER center for their exceptional work on68planning and execution of the data collection and annotation69exercise. We would also like to thank Rajarshi Singh from the70PAL Network and Uday Narayan Singh who helped us with71our sampling strategy. Lastly we would like to thank Dolly72Agarwal and everybody from the Tech team of Pratham Education Foundation for their work in making this project pos-73sible.74 75### Contact US76For the larger Dataset Access Contact satish.k@pratham.org , sravana.chandra@pratham.org77 