CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01armand0e /claude-fable-5-claude-code claude-fable-5 Agent Traces It's worth noting that our team was working with Glint-Research to collect as much fable data as possible. These are just the anonymized raw traces of both of our teams combined. This means that Glint-Research/Fable-5-traces was created from formatting and splitting up this same dataset. If you use one for your tune, don't use the other (it's the same exact data). For training on this dataset I recommend using the teich package to convert to openai… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-fable-5-claude-code.tabulartext-generationn<1K388 likes4.8k downloads14d agoHugging Face02armand0e /gpt-5.5-agentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. gpt 5.5 Agent Traces This directory contains raw agent trace files generated by teich. (I also dropped in some of my own personal traces) All assistant responses were generated by openai/gpt-5.5. JSONL files: 88 Training-ready tools A complete configured tools schema snapshot is embedded in the… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/gpt-5.5-agent.tabularn<1K18 likes1.1k downloads4mo agoHugging Face03marin-dna /gpn-star-p-uniform-v1-enhancer-arm-a marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses calibrated entropy from the primate… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/gpn-star-p-uniform-v1-enhancer-arm-a.tabular100M<n<1B0 likes873 downloads29d agoHugging Face04armand0e /kimi-k2.6-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 36 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-claude-code-traces.tabularn<1K4 likes832 downloads4mo agoHugging Face05armand0e /qwen3.7-max-pi-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Qwen3.7 Max Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by qwen/qwen3.7-max. JSONL files: 47 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README. Use it… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-max-pi-traces.tabulartext-generationn<1K96 likes775 downloads4mo agoHugging Face06armnet /armnet-demo-leaderboardtabularn<1K0 likes769 downloads20d agoHugging Face07marin-dna /phylop-uniform-v1-enhancer-arm-a marin-dna/phylop-uniform-v1-enhancer-arm-a Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the enhancer region cohort with all species scope and preserves source FASTA/2bit letter case. Non-human rows project only the central human nucleotide and extract the 255 bp target window centered on its unique mapped locus. Anchor eligibility uses the pipeline's pinned phyloP conservation… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/phylop-uniform-v1-enhancer-arm-a.tabular10M<n<100M0 likes600 downloads27d agoHugging Face08armand0e /minimax-m3-claude-code-tracesThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Minimax M3 Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by minimax/minimax-m3. JSONL files: 31 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m3-claude-code-traces.tabulartext-generationn<1K13 likes557 downloads4mo agoHugging Face09armand0e /teich-test-v1 hy3-preview coding agent traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by tencent/hy3-preview:free. Training-ready tools Use this tools payload when rendering converted examples through your training chat template. The same structure is emitted on each converted example as the tools field. [ { "type": "function", "function": { "name": "bash", "description": "Execute bash… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/teich-test-v1.tabularn<1K0 likes464 downloads5mo agoHugging Face10armand0e /minimax-m2.7-agent Agentic Training Traces This directory contains raw agent trace files generated by agentic-datagen. All assistant responses were generated by minimax/minimax-m2.7. Trace files: 20 Training-ready tools Use this tools payload when rendering converted examples through your training chat template. The same structure is emitted on each converted example as the tools field. [ { "type": "function", "function": { "name": "bash", "parameters": {… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/minimax-m2.7-agent.tabularn<1K0 likes268 downloads4mo agoHugging Face11ReasoningRegisters /arm1-eval-rollouts arm1 eval rollouts Eval rollouts (32 samples/problem) for ReasoningRegisters/arm1. Split folders named step{k}_{benchmark}[_variant]; JSONL per shard: prompt, generation, correctness. tabular10K<n<100K0 likes227 downloads19d agoHugging Face12armand0e /badlogicgames-pi-mono-opus-filteredFiltered version of badlogicgames/pi-mono - Only opus traces, dropped invalid sessions as well. All traces present are training safe and teich compatible tabularn<1K2 likes200 downloads4mo agoHugging Face13armand0e /kimi-k2.6-agentThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Kimi K2.6 Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by moonshotai/kimi-k2.6. JSONL files: 15 Training-ready tools A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of this README.… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/kimi-k2.6-agent.tabularn<1K2 likes196 downloads4mo agoHugging Face14ReasoningRegisters /arm3-eval-rollouts arm3 eval rollouts Eval rollouts (32 samples/problem) for ReasoningRegisters/arm3. Split folders named step{k}_{benchmark}[_variant]; JSONL per shard: prompt, generation, correctness. tabular10K<n<100K0 likes185 downloads19d agoHugging Face15procmarco /fpabl1-arm-b-fp-tokens-48k fpabl1-arm-b-fp-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-b-fp. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fp_en: 1,000,000,000 fp_ita:… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-b-fp-tokens-48k.tabularn<1K0 likes171 downloads3mo agoHugging Face16ReasoningRegisters /arm2-eval-rollouts arm2 eval rollouts Eval rollouts (32 samples/problem) for ReasoningRegisters/arm2. Split folders named step{k}_{benchmark}[_variant]; JSONL per shard: prompt, generation, correctness. tabular10K<n<100K0 likes171 downloads19d agoHugging Face17ReasoningRegisters /arm4-eval-rollouts arm4 eval rollouts Eval rollouts (32 samples/problem) for ReasoningRegisters/arm4. Split folders named step{k}_{benchmark}[_variant]; JSONL per shard: prompt, generation, correctness. tabular10K<n<100K0 likes154 downloads19d agoHugging Face18armand0e /qwen3.7-plus-claude-codeThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Qwen3.6 Plus Claude Code Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by qwen/qwen3.7-plus. JSONL files: 7 Format Each file is newline-delimited JSON representing a single captured agent session. The trace schema is designed for… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/qwen3.7-plus-claude-code.tabulartext-generationn<1K2 likes139 downloads4mo agoHugging Face19armand0e /claude-opus-4.8-pi-tracesMore expensive than anticpated so you only get 4 lol :P This dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. Claude Opus 4.8 Pi Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by anthropic/claude-opus-4.8. JSONL files: 4 Training-ready tools A complete configured tools schema snapshot is… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/claude-opus-4.8-pi-traces.tabulartext-generationn<1K8 likes137 downloads4mo agoHugging Face20armand0e /hermes-testThis dataset was generated using teich by TeichAI Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below. My Agent Traces This directory contains raw agent trace files generated by teich. All assistant responses were generated by nex-agi/nex-n2-pro:free. Sessions: 2 Training-ready tools Generated agent traces carry configured or recovered tool schemas so tools remain available for training even when a… See the full description on the dataset page: https://huggingface.co/datasets/armand0e/hermes-test.tabulartext-generationn<1K1 likes68 downloads3mo agoHugging Face21open-llm-leaderboard /RLHFlow__ArmoRM-Llama3-8B-v0.1-detailsgated Dataset Card for Evaluation run of RLHFlow/ArmoRM-Llama3-8B-v0.1 Dataset automatically created during the evaluation run of model RLHFlow/ArmoRM-Llama3-8B-v0.1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/RLHFlow__ArmoRM-Llama3-8B-v0.1-details.tabular10K<n<100K0 likes12 downloads2y agoHugging Face22procmarco /fpabl1-arm-a-web-tokens-48k fpabl1-arm-a-web-tokens-48k Pre-tokenized bins for the FinePhrase vs FineWeb ablation (fpabl1), arm fpabl1-a-web. Tokenizer: runs/mixed-tokenizer-48k/tokenizer (byte-BPE, 48k vocab; special ids bos=49119, eos=49120, pad=49121). Format: uint16 little-endian, 50 shards x 100,000,000 tokens = 5,000,000,000 tokens total. Boundary policy: source streams are already BOS/document/EOS packed; arm scheduler interleaves token chunks Source token composition: fw2_ita: 2,000,000,000… See the full description on the dataset page: https://huggingface.co/datasets/procmarco/fpabl1-arm-a-web-tokens-48k.tabularn<1K0 likes10 downloads3mo agoHugging Face23ehzoah /uf-ArmoRM-Llama3-8B-v0.1Relabled UltraFeedback dataset by ArmoRM-Llama3-8B-v0.1. tabular10K<n<100K0 likes9 downloads2y agoHugging Face24nishitneema /Magpie_coding_sfairXC_ArmoRM_top5_reward_differencetabularn<1K0 likes3 downloads2y agoHugging Face25dda71427 /robot_arm_6dof.jsontabularn<1K0 likes3 downloads9mo agoHugging Face26viswavi /wildchat_armormgatedtabular10K<n<100K0 likes2 downloads1y agoHugging Face27PJMixers /Intel_orca_dpo_pairs-ArmoRM-ReRanked-PreferenceShareGPTgatedtabular1K<n<10K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.