muril
Datasets
All datasets matching “muril”muril-nepali-gector-style-token-level-tag-for-ged
Nepali GEC (gector style) Token Tagging Dataset
This is a processed version of the sumitaryal/nepali_grammatical_error_correction dataset,
designed for training GEC-ToR-style sequence tagging models.
This dataset has been processed with a robust, multi-pass, content-aware alignment algorithm
to generate high-fidelity correction tags, including complex and adjacent SWAP operations.
Total Examples: 16,260,992
Training: 13,008,711
Validation: 2,439,231
test: 813,050… See the full description on the dataset page: https://huggingface.co/datasets/DipeshChaudhary/muril-nepali-gector-style-token-level-tag-for-ged.librispeech_ptmuril-research-data
Hindi MuRIL/GPT-2 Surprisal Research Data
Raw per-token surprisal embeddings and large derived result files from the Hindi word-order surprisal research project (Raja Linguistics Group).
Code
https://github.com/Raja-Linguistics-Group/hindi-muril-surprisal
The GitHub repo's scripts/download_data.py restores these files to their
original locations in the code repo; data_manifest.json there documents
what's here and where each file belongs.
Related… See the full description on the dataset page: https://huggingface.co/datasets/raja-linguistic-group/muril-research-data.voxforge_ptfootball-events-statsbomb360-la-liga-10-5kfootball-events-statsbomb360-la-liga-5-5k
