datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iNLTK_Sanskrit_Shlokas_DatasetRoundTripOCR-sanskritPost-OCR error correction dataset (train, test and validation set) for Sanskrit language generated using RoundTripOCR technique.
Code: https://github.com/harshvivek14/RoundTripOCR
Sanskriti
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
Link: arxiv.org/abs/2506.15355
Dataset Description
The SANSKRITI benchmark is the largest dataset created to evaluate Language Models' (LMs) comprehension and reasoning capabilities regarding the rich cultural diversity of India.It addresses the critical need for culturally-aware benchmarks, as the global effectiveness of LMs depends on their understanding of local… See the full description on the dataset page: https://huggingface.co/datasets/13ari/Sanskriti.Sanskriti
SANSKRITI: A Comprehensive Benchmark for Evaluating Language Models' Knowledge of Indian Culture
Link: arxiv.org/abs/2506.15355
Dataset Description
The SANSKRITI benchmark is the largest dataset created to evaluate Language Models' (LMs) comprehension and reasoning capabilities regarding the rich cultural diversity of India.It addresses the critical need for culturally-aware benchmarks, as the global effectiveness of LMs depends on their understanding of local… See the full description on the dataset page: https://huggingface.co/datasets/nishtharajput/Sanskriti.Sanskrit-Text-Summaryfine_tune_actual_dataxnli2.0_train_sanskritsanskritSanskritDatasetDCS_Sanskrit_Morphology_v1This dataset contains a dcs_output.csv file at the root.
xnli2.0_sanskritSanskritShlokaSanskritShloka3SanskritDatasetSanskrit_shlokasSanskritShloka2NIOS_Hindi_Sanskritfine_tuning_data_smallfine_tune_data_200sanskrit_sandhisfine_tune_2fine_tune_without_spSanskrit-Question-Answeringfine_tune_actual_5kfine_tune_without_prompt_5kfinal_data_3_samplefine_tune_llamadataset_sftdataset_sft_combinedfine_tune_llama_extended
