CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01schema-harness /arc-agi-3-schema-traces ARC-AGI-3 Schema Gameplay Trajectories This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free scoring utility. The trajectories are split evenly across two collections: gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories. claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5. Each trajectory directory includes run.json, a streamed events.jsonl event log, sanitized session data, snapshots, and the shareable text/image files produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.tabularn<1K38 likes1.4k downloads2mo agoHugging Face02JBrightmanAI /arc-agi-3-schema-traces ARC-AGI-3 Schema Gameplay Trajectories This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free scoring utility. The trajectories are split evenly across two collections: gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories. claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5. Each trajectory directory includes run.json, a streamed events.jsonl event log, sanitized session data, snapshots, and the shareable text/image files produced during… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/arc-agi-3-schema-traces.tabularn<1K0 likes407 downloads2mo agoHugging Face03guanning /arc-agi-3-schema-traces-gpt56gated ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol This release contains every gpt-5.6-sol gameplay trajectory produced on our cluster with the world_model_v5 agent harness — 100 runs across the 25 public ARC-AGI-3 games — plus a dependency-free scoring utility. It is the GPT-5.6 Sol member of a family built by the same harness and the same sanitizer, so trajectories can be compared game by game: arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.tabularreinforcement-learningn<1K0 likes81 downloads4d agoHugging Face04guanning /arc-agi-3-schema-traces-opus48gated ARC-AGI-3 Schema Gameplay Trajectories — Claude Opus 4.8 This release contains the best claude-opus-4-8 / max trajectory for each of the 25 public ARC-AGI-3 games, plus a dependency-free scoring utility. It is the Opus 4.8 counterpart of arc-agi-3-schema-traces-fable5, produced by the same agent harness (world_model_v5) and the same sanitizer, so the two collections can be compared game by game. Each trajectory directory includes run.json, a streamed events.jsonl event log… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-opus48.tabularreinforcement-learningn<1K0 likes47 downloads4d agoHugging Face05guanning /arc-agi-3-schema-traces-gpt56-xhighgated ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol (xhigh) The best gpt-5.6-sol trajectory at xhigh reasoning effort for each of the 25 public ARC-AGI-3 games, produced with the world_model_v5 agent harness. This release exists to make the cross-model comparison single-effort on all sides. Its siblings are each one model at one effort, but the gpt-5.6-sol collection in arc-agi-3-schema-gameplay is a mix of xhigh and max (16 games + 9 games), so it is not directly comparable to… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56-xhigh.tabularreinforcement-learningn<1K0 likes32 downloads2d agoHugging Face06npc-worldwide /npcsh-traces enpisi-coder RL dataset Judge-rated npcsh agent traces and derived RL training data for the enpisi-coder model family. Produced by scripts/rate_traces.py (LLM-as-judge) and scripts/analyze_ratings.py; built into SFT/DPO/GRPO/PPO splits by scripts/train_from_csv.py. Splits Split Rows Description rated_traces 38388 Per-trace judge scores (correctness, tool_selection, efficiency, clarity, partial_credit, composite) tasks 100 Benchmark task definitions… See the full description on the dataset page: https://huggingface.co/datasets/npc-worldwide/npcsh-traces.tabular100K<n<1M0 likes21 downloads2mo agoHugging Face07niranjanh123 /chatgpt_filtered_sft_traces_context_awaretabularreinforcement-learningn<1K0 likes16 downloads2mo agoHugging Face08niranjanh123 /sonnet_filtered_sft_traces_simplified_reasoningtabularreinforcement-learningn<1K0 likes11 downloads2mo agoHugging Face09niranjanh123 /gemini_filtered_sft_traces_simplified_reasoningtabularreinforcement-learningn<1K0 likes10 downloads2mo agoHugging Face10niranjanh123 /chatgpt_filtered_sft_traces_simplified_reasoningtabularreinforcement-learningn<1K0 likes10 downloads2mo agoHugging Face11sbhambr1 /cotempqa_for_sft_r1_set_custom_incorrect_tracestabular1K<n<10K0 likes8 downloads1y agoHugging Face12sbhambr1 /cotempqa_for_sft_r1_set_custom_tracestabular1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.