datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
codette-training-data
Codette Training Data — STaR Study Datasets
Self-generated reasoning datasets from the Codette STaR study. All chains
were generated by the Codette newton adapter (Llama 3.1 8B, OpenVINO INT4)
in reason mode and filtered as described.
File
Chains
Source
Filter
newton_star.jsonl
500
ARC-Challenge/OpenBookQA/SciQ train
keep-correct, ≥40 reasoning words (81% yield)
newton_star_hard.jsonl
350
MMLU-Pro STEM
keep-correct, ≥50 words (61% yield)
newton_star_rational.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette-training-data.codette_training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/Raiff1982/codette_training.Codettesspecialcodettefloodresponse
