FineEnvs/data-agent-sft
π οΈ Data Agent β SFT 4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool, reason, compute, and write the answer. Every one of these solved its task and passed a deterministic grader β so you're fine-tuning on demonstrations that are known to be correct, not just plausible. Drop-in ready for TRL: conversational messages + tools. Where it comesβ¦ See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-sft.
π οΈ Data Agent β SFT
4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool, reason, compute, and write the answer. Every one of these solved its task and passed a deterministic grader β so you're fine-tuning on demonstrations that are known to be correct, not just plausible.
Drop-in ready for TRL: conversational messages + tools.
Where it comes from
These are real agent rollouts on the Data Agent tasks, which were themselves built from the **jupyter-agent dataset** (data-science notebooks over Kaggle datasets). We kept only trajectories that reached the correct answer under deterministic grading (reward = 1.0) β one clean demonstration per task.
What's inside
- 4,677 correct trajectories β one per task
- Difficulty β easy 1,402 Β· medium 2,640 Β· hard 635
- One tool throughout:
bash(shell command execution)
What's in a row
- `messages` β the full conversation in OpenAI/TRL chat format:
systemβuser(the task) βassistant(reasoning +tool_calls) βtool(command output) β β¦ β finalassistantanswer. Tool-callargumentsare JSON objects;toolmessages carry the toolname. - `tools` β the
bashtool's JSON schema (rendered byapply_chat_template(..., tools=...)) task_id,difficulty(1β5),difficulty_tier,n_turns,source_agent
Fine-tune with TRL
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
ds = load_dataset("HuggingEnvs/data-agent-sft", split="train")
trainer = SFTTrainer(
model="Qwen/Qwen2.5-3B-Instruct", # any tool-capable chat template
train_dataset=ds, # messages + tools are picked up automatically
args=SFTConfig(assistant_only_loss=True, max_length=8192),
)
trainer.train()The messages + tools columns render through your model's chat template, and assistant_only_loss=True trains on the assistant's tokens only β no dataset wrangling needed.
