datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca_data_cleaned_bhojpuri
Dataset Card for Dataset Name
This repository contains a translated version of the Alpaca-Cleaned dataset, originally provided by Yahma on Hugging Face. The dataset has been translated into Bhojpuri, a language spoken in the northern-eastern part of India and the Terai region of Nepal.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
The Alpaca-Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/SatyamDev/alpaca_data_cleaned_bhojpuri.Bhojpuri-Behavioral-Corpus-8K
🚀 Bhojpuri Behavioral Corpus (Phase 2: Engineered Refinement)
⚠️ NOTICE: Phase 2 Refinement
This repository contains the Phase 2 Engineered Refinement. This Phase 2 is automatically refined specifically to prevent class collapse during fine-tuning.
📌 Executive Summary
The Bhojpuri Behavioral Corpus (Phase 2) is a 68,822-row, rigidly balanced dataset engineered to solve the inherent instability of low-resource language fine-tuning. Moving beyond noisy… See the full description on the dataset page: https://huggingface.co/datasets/abhiprd2000/Bhojpuri-Behavioral-Corpus-8K.xnli2.0_train_bhojpurixnli2.0_bhojpurikreol-bhojpuri-ratings
