CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Open-SWE-Traces Open-SWE-Traces: Advancing Distillation for Software Engineering Agents 🚨 What's New [09/26] Release v1.2: Added new agent trajectories generated by Qwen3.8-27B for mini-swe-agent. Trajectories for OpenCode and Claude Code harnesses will be released soon. [08/26] Release v1.1: Added new agent trajectories generated by DeepSeek-V4-Flash and Qwen3.6-27B across OpenHands, SWE-agent, and mini-swe-agent harnesses. [06/21] Release v1.0: Released 207k agent… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Open-SWE-Traces.text100K<n<1M131 likes32k downloads4d agoHugging Face02dacorvo /funes-nvidia-Open-SWE-Traces Funes recall store — NVIDIA Open-SWE-Traces (resolved) A funes recall store built by indexing the resolved==1 trajectories of nvidia/Open-SWE-Traces (65244 sessions, across both harnesses — SWE-agent and OpenHands — and both models, Minimax-M2.5 and Qwen3.5-122B). What this is This is not a raw trace dataset — it is a pre-built funes index: the source trajectories chunked into content blocks and embedded, stored as a Lance table (chunks.lance). Source… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/funes-nvidia-Open-SWE-Traces.10M<n<100M0 likes4.5k downloads2mo agoHugging Face03Jayfarei /opentraces-capsules0 likes1.9k downloads4mo agoHugging Face04genlabs /OpenTracesgatedtext-generation0 likes1.7k downloads3mo agoHugging Face05open-agent-leaderboard /traces0 likes1.2k downloads4mo agoHugging Face06open-athena /glm52-datagen-r11-106-instruction-following-citation-tracestext1K<n<10K1 likes845 downloads1mo agoHugging Face07open-athena /llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/llm-verifier-freelancer-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes307 downloads3mo agoHugging Face08open-athena /glm52-datagen-r11-100-agentic-function-calling-pivot-v2-tracestext1K<n<10K1 likes305 downloads1mo agoHugging Face09open-athena /nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-instruction-following-structured-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes303 downloads2mo agoHugging Face10open-athena /glm52-datagen-r11-safety-tracestext10K<n<100K1 likes302 downloads2mo agoHugging Face11open-athena /glm52-datagen-r11-16-nemotron-instruction-tracestext1K<n<10K0 likes279 downloads2mo agoHugging Face12open-athena /nemotron-gym-if-v2-qwen3.5-122b-32k-tracestext10K<n<100K0 likes277 downloads3mo agoHugging Face13open-athena /nemotron-math-oracle-filtered-qwen3.5-122b-32k-tracestext10K<n<100K0 likes258 downloads3mo agoHugging Face14kshitijthakkar /Open-SWE-Traces Open-SWE-Traces: Advancing Distillation for Software Engineering Agents Data Overview Open-SWE-Traces is an agentic instruction tuning dataset designed to advance the capabilities of LLMs in software engineering. This dataset comprises 200k+ agent trajectories collected using the SWE-agent and OpenHands framework. The trajectories were synthesized using Minimax-M2.5 (with thinking) and Qwen3.5-122B-A10B (without thinking) and specifically curated for supervised… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/Open-SWE-Traces.text100K<n<1M0 likes257 downloads2mo agoHugging Face15open-athena /stackexchange-tezos-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/stackexchange-tezos-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes254 downloads29d agoHugging Face16open-athena /nemotron-gym-identity-following-v2-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-identity-following-v2-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes244 downloads2mo agoHugging Face17open-athena /stackexchange-overflow-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/stackexchange-overflow-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes235 downloads27d agoHugging Face18open-athena /exp_rpt_pymethods2test-large-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_pymethods2test-large-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes234 downloads2mo agoHugging Face19OpenTrace /WorldTrace WorldTrace Dataset 🗺️ Overview WorldTrace is a large-scale, high-quality, globally covering GPS trajectory dataset. Trajectory data provides an important data source for understanding human mobility patterns and transforming urban intelligence. However, existing trajectory modeling methods have limitations in terms of task specificity, regional dependency, and data sensitivity. The construction of the WorldTrace dataset aims to address these challenges by… See the full description on the dataset page: https://huggingface.co/datasets/OpenTrace/WorldTrace.1M<n<10M9 likes229 downloads1y agoHugging Face20open-athena /glm52-datagen-r11-02-codecontests-tracestext1K<n<10K0 likes229 downloads2mo agoHugging Face21open-athena /nemotron-gym-instruction-following-structured-minimax-m27-131k-tracestext1K<n<10K0 likes228 downloads4mo agoHugging Face22open-athena /glm52-datagen-r11-42-ghactions-retry-r2-tracestext1K<n<10K0 likes225 downloads2mo agoHugging Face23open-athena /nemotron-gym-identity-following-v2-qwen3.5-122b-32k-tracestext1K<n<10K0 likes213 downloads3mo agoHugging Face24open-athena /stackexchange-superuser-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/stackexchange-superuser-sandboxes-verified-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes211 downloads1mo agoHugging Face25open-athena /nemotron-gym-agent-calendar-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-agent-calendar-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes196 downloads2mo agoHugging Face26open-athena /nemotron-gym-competitive-coding-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-competitive-coding-qwen3.5-122b-131k-opencode-traces.text10K<n<100K0 likes196 downloads2mo agoHugging Face27open-athena /exp_rle_adversarial-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rle_adversarial-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes195 downloads2mo agoHugging Face28open-athena /exp_rpt_pr-qwen3.5-122b-131k-opencode-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_pr-qwen3.5-122b-131k-opencode-traces.text1K<n<10K0 likes195 downloads2mo agoHugging Face29open-athena /selfinstruct-naive-sandboxes-2-verified-qwen3.5-122b-131k-opencode-tracestext1K<n<10K0 likes193 downloads2mo agoHugging Face30open-athena /glm52-datagen-r11-15-ghactions-tracestext10K<n<100K0 likes183 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.