datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-corpus
Filtered Multilingual Corpus (EN-HI-PA)
Dataset Description
A cleaned and balanced multilingual corpus extracted from the Samanantar dataset, containing parallel sentences in English, Hindi, and Punjabi.
Languages
The dataset contains text in three languages:
English (en)
Hindi (hi)
Punjabi (pa)
Dataset Summary
This dataset is a filtered subset of the Samanantar parallel corpus, specifically:
150,000 English sentences (from EN-HI and EN-PA… See the full description on the dataset page: https://huggingface.co/datasets/PredictiveManish/multilingual-corpus.chatalpaca-multiturn-predictive-state5
ChatAlpaca Multiturn Predictive State 5
Predictive-state supervision for Samantha multiturn latent-state pretraining with current-observation shortcut probes removed.
This version is derived from BRlkl/chatalpaca-multiturn-predictive-state4.
The main filter removes probes whose answer is recoverable from the visible current observation:
previous assistant message
current user message
That makes the remaining supervision much closer to the actual objective for the multiturn… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-predictive-state5.chatalpaca-multiturn-predictive-state4
ChatAlpaca Multiturn Predictive State 4
Balanced predictive-state supervision for Samantha multiturn latent-state pretraining.
This version is derived from BRlkl/chatalpaca-multiturn-predictive-state3 with shortcut probes removed.
The main removed failure mode is the overrepresented first-user-request probe whose target is often the full user prompt.
Audit
rows: 7816
probes: 15034
estimated expanded probe branches: 24272
max family share: 0.1995
answer equals… See the full description on the dataset page: https://huggingface.co/datasets/BRlkl/chatalpaca-multiturn-predictive-state4.
