CoolFace
Datasetpublic

AdithyaSK/data_agent_harbor_eval

data_agent_harbor_eval 144 deterministic data-analysis tasks for agent RL (validation split). Each task gives an agent a Kaggle dataset and a question; the answer is graded deterministically (exact -> numeric tolerance -> list/percent normalization -> symbolic, no LLM judge). Difficulty tiers: {'hard': 54, 'easy': 52, 'medium': 38}. Environments build from base image savatar101/env-data-agent-train:base. Format Harbor task suite: tasks/<id>/ (task.toml… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_harbor_eval.

sourceHugging Facemitupdated 22d agoView on Hugging Face
0likes739downloads
Dataset Card

dataagentharbor_eval

144 deterministic data-analysis tasks for agent RL (validation split). Each task gives an agent a Kaggle dataset and a question; the answer is graded deterministically (exact -> numeric tolerance -> list/percent normalization -> symbolic, no LLM judge).

Difficulty tiers: {'hard': 54, 'easy': 52, 'medium': 38}. Environments build from base image savatar101/env-data-agent-train:base.

Format

Harbor task suite: tasks/<id>/ (task.toml, instruction.md, environment/, tests/), registry.json, manifest.parquet.