code-agent
Fireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-i1-GGUFLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2-i1-GGUFqwen2.5-coder-7b-agent-ggufLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v3-i1-GGUFFireball-Meta-Llama-3.2-8B-Instruct-agent-003-128k-code-DPO-GGUFFireball-Meta-Llama-3.1-8B-Instruct-Agent-0.003-128K-code-ds-auto-GGUFLFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2-GGUFLlama-3.2-3B-Agent007-Coder-GGUF
da-code-evaluation-resultsTau2-Bench-Airline-With-Code-Agents
Dataset Card for a Code Agent Version of Tau Bench 2 Airline
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between an code agent and AI assistant. The dataset is based on the Airline environment from Tau^2 Bench and contains traces from both the original version and a version made at Snorkel AI using code agents to solve the same tasks (indicator in the version field; details below).
Curated by: Snorkel AI… See the full description on the dataset page: https://huggingface.co/datasets/snorkelai/Tau2-Bench-Airline-With-Code-Agents.ai-code-generation-swe-agents-2026
💻 AI Code Generation, SWE Agents & Program Synthesis Dataset (2026 Edition)
A structured research dataset featuring 3,181 domain-verified research papers and 771 official code repositories focused on Autonomous Software Engineering Agents (SWE-bench), Program Synthesis, DeepSeek-Coder-V2, Qwen2.5-Coder, Test-Driven Code Repair, Self-Healing Software, AST Semantic Modeling, and Formal Logic Verification (2023–2026).
Built with Universal Scientific Engine V17.1 Gold, providing 47… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/ai-code-generation-swe-agents-2026.works_on_my_agent_code
Track B Phase 3 Submission
Team: Works on my agent
This archive contains the runnable submission for Track B Phase 3.
Environment
Python 3.11 is recommended for the inference runner:
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
Our local validation environment used Huawei Ascend 910B hardware.
The runner does not require internet access at runtime. It connects only to the local vLLM… See the full description on the dataset page: https://huggingface.co/datasets/jinao/works_on_my_agent_code.agentic_code_dataset_22Dataset: 22 Real Claude Code Sessions
To validate Suffix Decoding's applicability in Agentic Coding scenarios, we collected 22 complete Claude Code session recordings.
Dataset Overview
Metric
Value
Collection date
December 2025
Total sessions
22
Total conversation turns
17,487
Total runtime
50 hours
Total input tokens
6,996,619
Total output tokens
6,094,906
Session Scale Distribution
Statistic
Min
Max
Average
Conversation turns
273… See the full description on the dataset page: https://huggingface.co/datasets/novita/agentic_code_dataset_22.agent-code-rl-artifacts
Agent Code RL Artifacts
Recovered process data from a code-generation Agent project covering SFT,
Monte Carlo rollout, process reward modeling, and veRL GRPO. This repository
contains benchmark-derived records and AI-generated content; it is not a
human-authored-only dataset.
Related SFT adapter:
keryszhan/qwen2.5-coder-7b-code-plan-sft.
Data stages
Config
Purpose
Important boundary
splits
Canonical HumanEval/MBPP-derived task splits
grpo_evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/keryszhan/agent-code-rl-artifacts.
