CoolFace
Datasetpublic

samarthramesh/gaperon-distill

gaperon-distill Reformatted, lean parquet build of the Gaperon mmBERT quality-distillation training data. One config per language (english, hindi, tamil), each with train / validation / test splits preserved exactly from the original make_splits partition (seed 42, 80/10/10). Built for fast loading on a cluster with no persistent storage. from datasets import load_dataset ds = load_dataset("samarthramesh/gaperon-distill", "hindi", split="train") Columns… See the full description on the dataset page: https://huggingface.co/datasets/samarthramesh/gaperon-distill.

sourceHugging Faceupdated 3mo agoView on Hugging Face
1likes113downloads
settings

This repository belongs to samarthramesh on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegaperon-distill
visibilitypublic
licencenot set
gatedno
ownersamarthramesh
Account settings
samarthramesh/gaperon-distill · CoolFace