CoolFace
Datasetpublic

schneiderkamplab/dfm10-openstax-open-chats

dfm10-openstax-open-chats Audited student inquiry conversations grounded in 61 immutable historical OpenStax CC BY 4.0 books. Contents Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz Schema: chat messages, optional condition and tools, plus provenance Shards: 1 Rows: 158,605 Category: Grounded English textbook student chats Upstream material OpenStax official immutable CC BY 4.0 source books Processing Eight… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm10-openstax-open-chats.

sourceHugging Facecc-by-4.0updated 27d agoView on Hugging Face
0likes47downloads
Dataset Card

dfm10-openstax-open-chats

Audited student inquiry conversations grounded in 61 immutable historical OpenStax CC BY 4.0 books.

Contents

  • —Format: gzip-compressed JSON Lines under data/train-*.jsonl.gz
  • —Schema: chat messages, optional condition and tools, plus provenance
  • —Shards: 1
  • —Rows: 158,605
  • —Category: Grounded English textbook student chats

Upstream material

  • —OpenStax official immutable CC BY 4.0 source books

Processing

Eight substantive pedagogical lenses are generated per verified passage; every assistant turn and complete conversation pass a Gemma 4 31B grounding audit.

Selection policy: all rows in the packaged source artifact.

Every packaged row is taken from the accepted source tree identified in the package manifest. Tokenized arrays and epoch sampling indices are not included; export staging alone does not imply inclusion in a sampled training union.

License and release review

The source books and this packaged derivative are distributed under Creative Commons Attribution 4.0. Preserve row-level OpenStax attribution. Preserve book title, immutable revision, source URL, artifact hash, and OpenStax attribution.

Validate

bash
python recreate_dataset.py