CoolFace
Datasetpublic

FineEnvs/data-agent-sft

πŸ› οΈ Data Agent β€” SFT 4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool, reason, compute, and write the answer. Every one of these solved its task and passed a deterministic grader β€” so you're fine-tuning on demonstrations that are known to be correct, not just plausible. Drop-in ready for TRL: conversational messages + tools. Where it comes… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-sft.

sourceHugging Facemitupdated 22d agoView on Hugging Face
0likes115downloads
Dataset Card

πŸ› οΈ Data Agent β€” SFT

4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool, reason, compute, and write the answer. Every one of these solved its task and passed a deterministic grader β€” so you're fine-tuning on demonstrations that are known to be correct, not just plausible.

Drop-in ready for TRL: conversational messages + tools.

Where it comes from

These are real agent rollouts on the Data Agent tasks, which were themselves built from the **jupyter-agent dataset** (data-science notebooks over Kaggle datasets). We kept only trajectories that reached the correct answer under deterministic grading (reward = 1.0) β€” one clean demonstration per task.

What's inside

  • β€”4,677 correct trajectories β€” one per task
  • β€”Difficulty β€” easy 1,402 Β· medium 2,640 Β· hard 635
  • β€”One tool throughout: bash (shell command execution)

What's in a row

  • β€”`messages` β€” the full conversation in OpenAI/TRL chat format: system β†’ user (the task) β†’ assistant (reasoning + tool_calls) β†’ tool (command output) β†’ … β†’ final assistant answer. Tool-call arguments are JSON objects; tool messages carry the tool name.
  • β€”`tools` β€” the bash tool's JSON schema (rendered by apply_chat_template(..., tools=...))
  • β€”task_id, difficulty (1–5), difficulty_tier, n_turns, source_agent

Fine-tune with TRL

python
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig

ds = load_dataset("HuggingEnvs/data-agent-sft", split="train")

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-3B-Instruct",     # any tool-capable chat template
    train_dataset=ds,                      # messages + tools are picked up automatically
    args=SFTConfig(assistant_only_loss=True, max_length=8192),
)
trainer.train()

The messages + tools columns render through your model's chat template, and assistant_only_loss=True trains on the assistant's tokens only β€” no dataset wrangling needed.