CoolFace
Datasetpublic

japhba/loracles-fineweb-multidoc-qa-1704

loracles-finetune-gemini-3.1-flash-lite Synthetic Loracle supervision data generated from FineWeb. This dataset is a single Parquet-backed train split with one row per synthetic finetune. This upload is a partial snapshot of a larger run. Run summary source dataset: HuggingFaceFW/fineweb / sample-10BT / train sampled docs: 15000 synthetic finetunes: 3009 generated finetunes uploaded: 2018 generator backend: openrouter generator model:… See the full description on the dataset page: https://huggingface.co/datasets/japhba/loracles-fineweb-multidoc-qa-1704.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes39downloads
Dataset Card

loracles-finetune-gemini-3.1-flash-lite

Synthetic Loracle supervision data generated from FineWeb.

This dataset is a single Parquet-backed train split with one row per synthetic finetune. This upload is a partial snapshot of a larger run.

Run summary

  • —source dataset: HuggingFaceFW/fineweb / sample-10BT / train
  • —sampled docs: 15000
  • —synthetic finetunes: 3009
  • —generated finetunes uploaded: 2018
  • —generator backend: openrouter
  • —generator model: google/gemini-3.1-flash-lite-preview
  • —max docs per finetune: 20
  • —max token budget per finetune: 10000
  • —questions per finetune: 10

Table schema

  • —one row per synthetic finetune
  • —nested questions field with 10 ordered question-answer pairs
  • —finetune metadata merged in from synthetic_models.jsonl
  • —stored on the Hub as Parquet

Local artifacts

  • —run directory: /ceph/scratch/jbauer/loracle/fineweb_openrouter_flash31lite_large_20260417
  • —sample manifest: sample_manifest.json
  • —question-group shards uploaded: 3
  • —generation manifests uploaded: 2
  • —auxiliary files kept as repo artifacts: synthetic_models.jsonl, sampled_docs.jsonl, question_failures.shard-*.jsonl