CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YYYYYYibo /alfworld-experimenter-gpt5mini-sft-1k ALFWorld Experimenter GPT-5 mini SFT 1K This dataset contains 1,000 blind GPT-5 mini reasoning demonstrations for an ALFWorld expert-prefix selection task. The intended use is to give a 7B experimenter model a structured reasoning warm start before reinforcement learning, not to treat GPT-5 mini's selected depths as ground-truth labels. Task For each ALFWorld task, the experimenter receives eight failed trajectories from a frozen Qwen2.5-7B-Instruct actor and one… See the full description on the dataset page: https://huggingface.co/datasets/YYYYYYibo/alfworld-experimenter-gpt5mini-sft-1k.tabulartext-generation1K<n<10K0 likes85 downloads19d agoHugging Face02jash404 /emergent-misalignment-experiment-1-data Emergent Misalignment Experiment 1 Data Artifacts Curated SFT data and diagnostics for an awareness-stratified code experiment on emergent misalignment. This artifact contains the exact trainable JSONL branches used for the reported n=1000 and n=3452 runs, plus the small manifests and balance summaries needed to audit the data mixture. The paired model adapters are available at jash404/emergent-misalignment-experiment-1-adapters. The source code and reports are in… See the full description on the dataset page: https://huggingface.co/datasets/jash404/emergent-misalignment-experiment-1-data.tabulartext-generationn<1K0 likes80 downloads4mo agoHugging Face03Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads25d agoHugging Face04ml-intern-explorers /experiment-1-oracles Experiment 1 Oracles Corpus Dataset Summary Dataset containing synthetic ground-truth oracle responses and intermediate reasoning traces for multi-step AI experiment evaluation. tabulartext-generation100K<n<1M1 likes11 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.