CoolFace
Datasetpublic

jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b

Independent synthetic biographies, role B This public dataset is the authenticated 3,550-person role-B prefix used by the scratch GPT-2-medium A/B experiment. It contains 32 independently generated biography views per person (113,600 rows) and a separate canonical four-question QA bundle per person (14,200 rows). The biographies were generated with the pinned openai/gpt-oss-120b v37 workflow. The biographies configuration exposes split train; the qa configuration exposes split… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes20downloads
Dataset Card

Independent synthetic biographies, role B

This public dataset is the authenticated 3,550-person role-B prefix used by the scratch GPT-2-medium A/B experiment. It contains 32 independently generated biography views per person (113,600 rows) and a separate canonical four-question QA bundle per person (14,200 rows).

The biographies were generated with the pinned openai/gpt-oss-120b v37 workflow. The biographies configuration exposes split train; the qa configuration exposes split test. QA rows were not used for biography SFT.

The two published roles have disjoint person names. Exact source, population, pair, file, and exclusion provenance is under provenance/. The authenticated GPT-2 training token cache is intentionally not public because it is a derived training optimization rather than dataset content.

Repository: jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b