CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ar0cket1 /hintedselfteacher-nemotron-math-v2-AoPS hintedselfteacher-nemotron-math-v2-AoPS This dataset contains a training-ready hinted self-teacher split derived from the AoPS split of nvidia/Nemotron-Math-v2. The source problems were filtered to the AoPS split with the medium/notool solve rate between 2 and 6. Hints were generated with GPT-5.5 medium using an h17_nt hint-generation prompt. This hint type was close to the best hint type found after doing hint mutations, based on qualitative analysis of token-level hinted… See the full description on the dataset page: https://huggingface.co/datasets/ar0cket1/hintedselfteacher-nemotron-math-v2-AoPS.tabulartext-generation10K<n<100K1 likes79 downloads3mo agoHugging Face02ceselder /cot-oracle-eval-hinted-mcq CoT Oracle Eval: hinted_mcq GSM8K problems as 4-choice MCQ with hints. 50/50 right/wrong hints, varying subtlety. Source: openai/gsm8k test. Part of the CoT Oracle Evals collection. Schema Field Description eval_name Eval identifier example_id Unique example ID clean_prompt Prompt without nudge/manipulation test_prompt Prompt with nudge/manipulation correct_answer Ground truth answer nudge_answer Answer the nudge pushes toward… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-oracle-eval-hinted-mcq.textn<1K0 likes47 downloads7mo agoHugging Face03jkarns /hinted_mbpp_llama2_7B_chat Dataset Card for "hinted_mbpp_llama2_7B_chat" More Information needed textn<1K0 likes42 downloads3y agoHugging Face04UnfaithRL /mmlu_hinted_questions MMLU Hinted Questions Dataset Description This dataset contains multiple-choice questions derived from MMLU and augmented with misleading hints. The misleading hints are intentionally designed to point to an incorrect answer. The dataset was developed as part of the UnfaithRL project, which studies cue-following and unfaithful reasoning under reinforcement learning with verifiable rewards. Specifically, it was used to investigate whether language models follow… See the full description on the dataset page: https://huggingface.co/datasets/UnfaithRL/mmlu_hinted_questions.tabularquestion-answering10K<n<100K0 likes41 downloads3mo agoHugging Face05lucferon /mmlu_hinted_rollouts MMLU-with-hint faithfulness eval — flipped-to-hint rollouts (+ judge verdicts) Companion data for the blog post on side effects of CoT length penalties in RL (MATS sprint project). Model checkpoints: brikdavies/RL-length-penalty-checkpoints. Each row is one MMLU question (~5k question eval, hint placed mid-prompt) where the model flipped its answer to the hinted answer (unhinted_answer != hinted_answer and the hinted run's extracted answer equals the hint). Rows carry: the… See the full description on the dataset page: https://huggingface.co/datasets/lucferon/mmlu_hinted_rollouts.tabular100K<n<1M0 likes30 downloads2mo agoHugging Face06alonmiron /mmlu_hinted_huggingfaceThis is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.question-answering0 likes26 downloads2y agoHugging Face07alonmiron /hinted_mmlu_560This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.0 likes18 downloads2y agoHugging Face08ceselder /cot-oracle-eval-hinted-mcq-truthfulqatabularn<1K0 likes18 downloads7mo agoHugging Face09Shahradmz /education_qna_hinted_statictextn<1K0 likes15 downloads2y agoHugging Face10alonmiron /mmlu_1120_hintedThis is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.0 likes13 downloads2y agoHugging Face11alonmiron /mmlu_1120_hinted_testThis is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.0 likes13 downloads2y agoHugging Face12Shahradmz /education_qna_hintedtextn<1K0 likes10 downloads2y agoHugging Face13TAUR-dev /lmfd__hinted_method__gpt4omini Dataset card for lmfd__hinted_method__gpt4omini This dataset was made with Curator. Dataset details A sample from the dataset: { "question": "What is the solution to the long multiplication equation below?\n\n8274 x 3529\n\nThink step by step.", "solution": "29198946", "eval_prompt": "What is the solution to the long multiplication equation below?\n\n8274 x 3529\n\nThink step by step.\n\nBreak this question down step by step using the **distributive… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/lmfd__hinted_method__gpt4omini.text1K<n<10K0 likes8 downloads1y agoHugging Face14alonmiron /test_mmlu_hintedThis dataset contains a copy of the cais/mmlu HF dataset but without the auxiliary_train split that takes a long time to generate again each time when loading multiple subsets of the dataset. Please visit https://huggingface.co/datasets/cais/mmlu for more information on the MMLU dataset. question-answering0 likes6 downloads2y agoHugging Face15TAUR-dev /evals__lmfd__hinted_method__gpt4omini__samplestabularn<1K0 likes5 downloads1y agoHugging Face16TAUR-dev /convos__lmfd__hinted_method__gpt4ominitext1K<n<10K0 likes4 downloads1y agoHugging Face17TAUR-dev /lmfd__hinted_method_w_verification__gpt4omini Dataset card for lmfd__hinted_method_w_verification__gpt4omini This dataset was made with Curator. Dataset details A sample from the dataset: { "question": "What is the solution to the long multiplication equation below?\n\n8274 x 3529\n\nThink step by step.", "solution": "29198946", "eval_prompt": "What is the solution to the long multiplication equation below?\n\n8274 x 3529\n\nThink step by step.\n\nBreak this question down step by step using the… See the full description on the dataset page: https://huggingface.co/datasets/TAUR-dev/lmfd__hinted_method_w_verification__gpt4omini.textn<1K0 likes4 downloads1y agoHugging Face18Shahradmz /education_qna_hinted_qwen05textn<1K0 likes3 downloads1y agoHugging Face19TAUR-dev /evals__full_TAUR_dev__convos__lmfd__hinted_method__gpt4omini__samplestabularn<1K0 likes3 downloads1y agoHugging Face20TAUR-dev /evals__full_TAUR_dev__convos__lmfd__hinted_method__gpt4omini__resultstextn<1K0 likes3 downloads1y agoHugging Face21TAUR-dev /evals__lmfd__hinted_method__gpt4omini__resultstextn<1K0 likes3 downloads1y agoHugging Face22ml-in-mind /postgresql-hinted-jobceb0 likes2 downloads6mo agoHugging Face23alonmiron /mmlu_560_haiku_hintedtextn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.