datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linear_probingedge_probing_dep_ewt_line_by_lineProbingPanoia-v01
ProbingPanoia
Train your models to ask questions.
I had google/gemma-2-9b-it and Qwen/Qwen2.5-7B-Instruct generate clarifying questions to a given instruction, then I used LittleInstructionMaker to generate the answer, and lastly mistralai/Mistral-Nemo-Instruct-2407 to generate the final turn.
Prompt template
I am synthesising a new dataset to teach language models how to ask for more context. You are tasked with creating responses to instruction prompts asking for… See the full description on the dataset page: https://huggingface.co/datasets/trollek/ProbingPanoia-v01.Numeracy-ProbingThe arXiv data used in LLMs Know More About Numbers than They Can Say
(arXiv:2602.07812,
code).
The dataset is derived from
allenai/peS2o.
blocksworld-4-self-probing-parsed-bigblocksworld-4-self-probing-parsed-big-v2
