CoolFace
Datasetpublic

leonli66/stage3-real-expansion-agent-teacher-separated-pilot

Teacher-Separated Expansion Agent Pilot A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks. The teacher-only trajectory-generation system prompt is recorded in metadata/generation-manifest.json for auditability, but is absent from every saved training trajectory. Each final messages list begins with the real memory-wrapped task user message, followed by native assistant expand calls, exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.

sourceHugging Faceupdated 24d agoView on Hugging Face
0likes58downloads
Dataset Card

Teacher-Separated Expansion Agent Pilot

A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.

The teacher-only trajectory-generation system prompt is recorded in metadata/generation-manifest.json for auditability, but is absent from every saved training trajectory. Each final messages list begins with the real memory-wrapped task user message, followed by native assistant expand calls, exact tool results, and the final assistant answer. The canonical native tool schema remains in the row-level tools field.

  • —Attempts: 10
  • —Accepted train rows: 5
  • —Rejected audit rows: 5
  • —Acceptance rate: 50%
  • —Total native expansion calls: 14
  • —System messages in saved trajectories: 0
  • —Full source tasks: audit/source-tasks.jsonl