CoolFace
Datasetpublic

ceselder/loracle-pretrain-mix-oneq

loracle-pretrain-mix This is the oneq subsample: one randomly-selected QA row per organism_id (deterministic shuffle with seed=42, then drop_duplicates). Same 3 splits, same row schema, half the rows. Source: ceselder/loracle-pretrain-mix. Built for loracle-training scale ablations where we want each training step to expose the model to a fresh organism (no 2-QA-per-org redundancy). Split sizes: data/train.parquet: 25000 rows (25000 organisms) data/dpo_heldout.parquet: 250… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/loracle-pretrain-mix-oneq.

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes8downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ceselder/loracle-pretrain-mix-oneq · CoolFace