janani-rane/SiQuAD
Sinhala SQuAD Dataset This dataset is a translation of the SQuAD v1.0 dataset into Sinhala using the Google Translate API. It consists of 16,000 question-answer pairs, with 13,000 training pairs and 1,250 test/dev pairs. The dataset is cleaned and validated to ensure the quality of the translations. Dataset Details Size: 16,000 QA pairs Train: 13,000 pairs Test/Dev: 1,250 pairs Language: Sinhala Source: The dataset was derived from the original SQuAD v1.0… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/SiQuAD.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face