CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01homerquan /boardgamebench-answer-only BoardGameBench Answer-Only Reasoning Dataset This dataset contains 1,282,766 board-game reasoning examples generated from BoardGameBench, a benchmark and data-generation project for evaluating language models on structured board-game decision making. Each row asks a model to inspect a legal board position and return the best move. The format is intentionally simple: id,prompt,answer The answer field is the target move label, such as C4, f6, 11,7, or e2-d3. This makes the dataset… See the full description on the dataset page: https://huggingface.co/datasets/homerquan/boardgamebench-answer-only.text-generation1M<n<10M0 likes153 downloads5mo agoHugging Face02Xiaofeng77 /answer-only-gp-l-only-10k Debunk the Myth of SFT Generalization Dataset This dataset is associated with the paper "Debunk the Myth of SFT Generalization". The paper challenges the prevailing view that supervised fine-tuning (SFT) primarily memorizes training data and fails to generalize, in contrast to reinforcement learning (RL). It demonstrates that SFT can generalize as well as—or better than—RL when trained with appropriate data, achieved through prompt diversity and Chain-of-Thought (CoT) supervision on… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/answer-only-gp-l-only-10k.texttext-generation10K<n<100K0 likes52 downloads1y agoHugging Face03Xiaofeng77 /diverse-answer-only-gp-l-only-10k General Points Dataset from Debunk the Myth of SFT Generalization This dataset is part of the research presented in the paper Debunk the Myth of SFT Generalization. It contains data for the General Points decision-making benchmark, which is used to evaluate the generalization capabilities of Supervised Fine-Tuning (SFT) models against Reinforcement Learning (RL) baselines. The paper explores the impact of prompt diversity and Chain-of-Thought (CoT) supervision on SFT's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-answer-only-gp-l-only-10k.texttext-generation10K<n<100K0 likes22 downloads1y agoHugging Face04Xiaofeng77 /diverse-answer-only-sokoban Dataset from "Debunk the Myth of SFT Generalization" This dataset is associated with the research presented in the paper Debunk the Myth of SFT Generalization. The paper challenges the conventional wisdom that supervised fine-tuning (SFT) primarily memorizes training data and struggles with generalization, contrasting it with reinforcement learning (RL)'s perceived robustness. Through systematic evaluation on decision-making benchmarks such as Sokoban and General Points, the… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/diverse-answer-only-sokoban.texttext-generation1K<n<10K0 likes16 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.