janani-rane/SiQuAD
Sinhala SQuAD Dataset This dataset is a translation of the SQuAD v1.0 dataset into Sinhala using the Google Translate API. It consists of 16,000 question-answer pairs, with 13,000 training pairs and 1,250 test/dev pairs. The dataset is cleaned and validated to ensure the quality of the translations. Dataset Details Size: 16,000 QA pairs Train: 13,000 pairs Test/Dev: 1,250 pairs Language: Sinhala Source: The dataset was derived from the original SQuAD v1.0… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/SiQuAD.
This repository belongs to janani-rane on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
