japhba/loracles-fineweb-multidoc-qa-1704
loracles-finetune-gemini-3.1-flash-lite Synthetic Loracle supervision data generated from FineWeb. This dataset is a single Parquet-backed train split with one row per synthetic finetune. This upload is a partial snapshot of a larger run. Run summary source dataset: HuggingFaceFW/fineweb / sample-10BT / train sampled docs: 15000 synthetic finetunes: 3009 generated finetunes uploaded: 2018 generator backend: openrouter generator model:… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-fineweb-multidoc-qa-1704.
loracles-finetune-gemini-3.1-flash-lite
Synthetic Loracle supervision data generated from FineWeb.
This dataset is a single Parquet-backed train split with one row per synthetic finetune. This upload is a partial snapshot of a larger run.
Run summary
- source dataset:
HuggingFaceFW/fineweb/sample-10BT/train - sampled docs:
15000 - synthetic finetunes:
3009 - generated finetunes uploaded:
2018 - generator backend:
openrouter - generator model:
google/gemini-3.1-flash-lite-preview - max docs per finetune:
20 - max token budget per finetune:
10000 - questions per finetune:
10
Table schema
- one row per synthetic finetune
- nested
questionsfield with10ordered question-answer pairs - finetune metadata merged in from
synthetic_models.jsonl - stored on the Hub as Parquet
Local artifacts
- run directory:
/ceph/scratch/jbauer/loracle/fineweb_openrouter_flash31lite_large_20260417 - sample manifest:
sample_manifest.json - question-group shards uploaded:
3 - generation manifests uploaded:
2 - auxiliary files kept as repo artifacts:
synthetic_models.jsonl,sampled_docs.jsonl,question_failures.shard-*.jsonl
