CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dlab-spp /reflection-50m SPP Reflection 50M The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero — the production half-corpus run, and the dataset the released models were actually trained on. 🔬 Small sample (same format): dlab-spp/reflection-sample-2k 📉 Earlier 10M run: dlab-spp/reflection-10m 🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.tabulartext-generation10M<n<100M0 likes1.4k downloads1mo agoHugging Face02OALL /details_SenseLLM__ReflectionCoder-DS-33B Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-DS-33B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-DS-33B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-DS-33B.tabular100K<n<1M0 likes1.2k downloads2y agoHugging Face03dlab-spp /reflection-10m SPP Reflection 10M The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection. Each row pairs a pretraining document with a synthetic, value-laden reflection generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.tabulartext-generation1M<n<10M0 likes341 downloads1mo agoHugging Face04feedbackagent /reflection_eval_prompt1text100K<n<1M0 likes292 downloads2y agoHugging Face05OALL /details_terrycraddock__Reflection-Llama-3.1-8B Dataset Card for Evaluation run of terrycraddock/Reflection-Llama-3.1-8B Dataset automatically created during the evaluation run of model terrycraddock/Reflection-Llama-3.1-8B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_terrycraddock__Reflection-Llama-3.1-8B.tabular100K<n<1M0 likes258 downloads2y agoHugging Face06OALL /details_SenseLLM__ReflectionCoder-CL-34B Dataset Card for Evaluation run of SenseLLM/ReflectionCoder-CL-34B Dataset automatically created during the evaluation run of model SenseLLM/ReflectionCoder-CL-34B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_SenseLLM__ReflectionCoder-CL-34B.tabular100K<n<1M0 likes235 downloads2y agoHugging Face07feedbackagent /reflection_eval_prompt2text100K<n<1M0 likes176 downloads2y agoHugging Face08feedbackagent /test_reflection_eval_prompttext100K<n<1M2 likes153 downloads2y agoHugging Face09feedbackagent /llama3_8b_reflection2text100K<n<1M2 likes86 downloads2y agoHugging Face10feedbackagent /train_reflection_eval1text10K<n<100K0 likes54 downloads2y agoHugging Face11feedbackagent /reflection_n16text10K<n<100K0 likes53 downloads2y agoHugging Face12TAUR-dev /9_8_25__countdown_4arg__sft_data_mp_reflection_ckpt_chunk_8text1K<n<10K0 likes53 downloads1y agoHugging Face13Harshkmr /orca-math-word-reflection Dataset Card for "orca-math-word-reflection" Dataset Summary The Orca-Math Word Problems with Reflection dataset is an extension of subset of the original ORCA Math Word Problems 200k dataset. This new version introduces a "Thinking and Reflection" format designed to enhance problem-solving approaches by encouraging step-by-step thinking before producing a solution. In this dataset, each math word problem and its corresponding solution from the original dataset are… See the full description on the dataset page: https://huggingface.co/datasets/Harshkmr/orca-math-word-reflection.texttext-generation1K<n<10K9 likes52 downloads2y agoHugging Face14TAUR-dev /9_8_25__letter_countdown_4o__sft_data_mp_reflection_ckpt_chunk_5tabular1K<n<10K0 likes52 downloads1y agoHugging Face15gabrielmbmb /distilabel-reflection-tuning Dataset Card for distilabel-reflection-tuning This dataset has been created with distilabel. The pipeline script was uploaded to easily reproduce the dataset: reflection.py. It can be run directly using the CLI: distilabel pipeline run --script "https://huggingface.co/datasets/gabrielmbmb/distilabel-reflection-tuning/raw/main/reflection.py" Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/distilabel-reflection-tuning.textn<1K55 likes49 downloads2y agoHugging Face16feedbackagent /train_reflection_eval2_with_rewardstext100K<n<1M0 likes48 downloads2y agoHugging Face17feedbackagent /train_reflection_eval2_with_rewards2text100K<n<1M0 likes43 downloads2y agoHugging Face18dvilasuero /reflection-v1-exact-duplicatestext10K<n<100K0 likes37 downloads2y agoHugging Face19feedbackagent /test_reflection_eval_completions_6text10K<n<100K0 likes37 downloads2y agoHugging Face20mychen76 /synthetic-think-and-reflection_v1This synthetic dataset was generated to improvement model thinking and reflection in systematical way. Columns question answer_with_tags text (processed for Llama 3.1) Dataset DatasetDict({ train: Dataset({ features: ['question', 'answer_with_tags', 'text'], num_rows: 17129 }) test: Dataset({ features: ['question', 'answer_with_tags', 'text'], num_rows: 52 }) }) Credit inspired by mattshumer (sharedGPT dataset) and… See the full description on the dataset page: https://huggingface.co/datasets/mychen76/synthetic-think-and-reflection_v1.text10K<n<100K3 likes32 downloads2y agoHugging Face21TAUR-dev /skillfactory_sft_countdown_3arg_qrepeat1_reflections5_formats0C.-C.-C-IC.-CCtext10K<n<100K0 likes31 downloads1y agoHugging Face22cmcmaster /reflection-v1text10K<n<100K0 likes30 downloads2y agoHugging Face23gardner /reflection-v1-sharegpttext10K<n<100K0 likes29 downloads2y agoHugging Face24feedbackagent /test_reflection_eval_completion1_with_rewardstext10K<n<100K0 likes29 downloads2y agoHugging Face25TAUR-dev /9_8_25__countdown_3arg__sft_data_mp_reflection_ckpt_chunk_59text1K<n<10K0 likes29 downloads1y agoHugging Face26TAUR-dev /OT_Ref_NoV13Part___openthoughts__1600000_end2000000__reflection_chunk_3text10K<n<100K0 likes29 downloads10mo agoHugging Face27xDAN-datasets /Maggen-Reflection-3.1-70b-50k-filtered-scoredDatasetDict({ train: Dataset({ features: ['created', 'response', 'pre_query_template', 'instruction', 'gen_input_configs', 'gen_response_configs', 'raw_instruction', 'id', 'instruction_sanitize_class_num', 'scores', 'model_name'], num_rows: 36884 }) }) 每个唯一值的计数: scores [9.0] 9469 [7.0] 6224 [10.0] 6009 [6.0] 4003 [8.0] 3149 [5.0] 2578 [4.0] 2575 [3.0] 1566 [2.0] 1051 [1.0] 209 [] 51 tabular10K<n<100K1 likes28 downloads2y agoHugging Face28rumbleFTW /reflection-v0text1K<n<10K0 likes28 downloads2y agoHugging Face29dlab-spp /reflection-sample-2k SPP Reflection 2k Sample A 2,000-row sample (seed 42) of dlab-spp/reflection-10m, in the identical format, for quick inspection of the data from Synthetic Persona Pretraining (SPP): Alignment from Token Zero. 📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero 📦 Full dataset: dlab-spp/reflection-10m (~10M documents). Each row pairs a pretraining document with a synthetic, value-laden reflection (first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.tabulartext-generation1K<n<10K0 likes28 downloads1mo agoHugging Face30feedbackagent /test_reflection_eval_completion4_with_rewardstext10K<n<100K0 likes26 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.