datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PTCCThe Parallel Tunisian Constitution Corpus (PTCC) corpus is a corpus of 149 articles written in Modern Standard Arabic and Tunisian Arabic.
Tesseract was used to transform the constitution's pdf files into text files. Afterward, alignment of the parallel articles was achieved by a simple Python script.
More details can be found in: https://amr-keleg.github.io/projects/digitalizing_dialectal_arabic/
Sources:
Tunisian Arabic translation of the 2014 Tunisian Constitution
2014… See the full description on the dataset page: https://huggingface.co/datasets/AMR-KELEG/PTCC.tourism-wellness-datasetamrutaDBThis dataset clone from amruta.org for training LLM
Contact: hungbui@sahajayoga.edu.vn
By the grace of Our H.H. Shri Mataji Nirmala Devi
amrutaQAThis dataset crawl and re-organize from site: amruta.org, extracting questions and answers for fine-tuning tasks.
By the grace of H.H. Mother Shri Mataji Nirmala Devi
text_classificationwastewater_amrjudicialhelprams-scramble-AMRcnn-sent-amr
Amr for each sentence in cnn
AmritaDatasetWilliam-AMRamr_geopolitical_riskTextClassificationAmritaDataset
