s1hikha/premchand-148-stories-dataset
Premchand 148 Stories Dataset A curated Hindi NLP corpus built from 148 short stories by Munshi Premchand, processed into sentence-aware chunks for RAG, embedding, and research pipelines. [!NOTE] This work is for a PhD in University of Allahabad by Shikha Agrawal under the supervision of Dr. Pravin Kumar. Dataset Structure The repository contains the following data files optimized for direct usage in any retrieval or generation model: corpus.jsonl: The final… See the full description on the dataset page: https://huggingface.co/datasets/s1hikha/premchand-148-stories-dataset.
034
