CoolFace
Datasetpublic

samarthramesh/gaperon-distill

gaperon-distill Reformatted, lean parquet build of the Gaperon mmBERT quality-distillation training data. One config per language (english, hindi, tamil), each with train / validation / test splits preserved exactly from the original make_splits partition (seed 42, 80/10/10). Built for fast loading on a cluster with no persistent storage. from datasets import load_dataset ds = load_dataset("samarthramesh/gaperon-distill", "hindi", split="train") Columns… See the full description on the dataset page: https://huggingface.co/datasets/samarthramesh/gaperon-distill.

sourceHugging Faceupdated 3mo agoView on Hugging Face
1likes115downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
samarthramesh/gaperon-distill · CoolFace