CoolFace
Datasetpublic

apodex/FrontierChallenge

FrontierChallenge FrontierChallenge provides 97 scientific workflow tasks with plaintext English instructions, inputs, Harbor definitions, domain labels, and the redistributable open runtime image. Path Contents manifest.jsonl Dataset Viewer rows with task ID, taxonomy, difficulty, runtime, and instruction tasks/<task-id>/ instruction.md, task metadata, environment definition, and agent-visible inputs images/ Verified linux/amd64 Docker archive for the 81… See the full description on the dataset page: https://huggingface.co/datasets/apodex/FrontierChallenge.

sourceHugging Facecc-by-4.0updated 2d agoView on Hugging Face
9likes2.2kdownloads
Dataset Card

FrontierChallenge

FrontierChallenge provides 97 scientific workflow tasks with plaintext English instructions, inputs, Harbor definitions, domain labels, and the redistributable open runtime image.

PathContents
manifest.jsonlDataset Viewer rows with task ID, taxonomy, difficulty, runtime, and instruction
tasks/<task-id>/instruction.md, task metadata, environment definition, and agent-visible inputs
images/Verified linux/amd64 Docker archive for the 81 open-image tasks

This repository contains no graders, rubrics, fixtures, or reference outputs. Those are stored as encrypted archives in the separate `apodex/FrontierChallenge-reference` dataset.

Use the FrontierChallenge runtime to download both datasets, verify their shared registry, load the image archive, run Harbor, and score a task. The evaluator host needs Python 3.12+ for the pinned Harbor 0.20.0; task containers keep their own frozen Python versions.

bash
git clone https://github.com/ApodexAI/FrontierAgent.git
cd FrontierAgent/benchmarks/frontierchallenge
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
HF_TOKEN=hf_... ./scripts/setup.sh --track open
cp .env.example .env
# Fill model and judge credentials in .env before running.
./scripts/run_eval.sh \
  --agent claude-code --model <model> \
  --include-task-name task_011_cell_migration_wound_healing

The 16 task definitions that execute ORCA and their inputs are included normally. ORCA itself and an ORCA-configured image are never distributed; evaluators obtain ORCA from its official provider and follow the runtime repository's local-image tutorial.

Track membership follows what a task executes, not software named in supplied files. For example, task_098_orca_claisen_thermochemistry reads precomputed ORCA output without running ORCA, so it remains in the open track.

Verify a downloaded solve package with:

bash
python tools/verify_dataset.py

Metrics

Official Pass Rate uses a strict score threshold: a task passes only when evaluation_complete == 1 and valid task_score > 0.999. Native per-task thresholds do not determine this metric. Use unrounded scores: exactly 0.999 does not pass; 0.9991 and 1.0 pass.

  • —Pass Rate (%) = 100 × number of completed tasks with valid task_score > 0.999 / 97.
  • —Score (0–100) = 100 × sum of valid, completed task_score values / 97.
  • —Partial credit remains the native rubric score normalized to [0, 1]. Missing tasks, invalid scores, and incomplete evaluations contribute zero to the fixed denominator and must be reported separately; incomplete runs are not final benchmark results.
  • —Use one predeclared attempt per task. A subset may use its predeclared task count as denominator, but must be labeled as a subset, not the 97-task result.

There is only one pass field, passed, using the same strict threshold in reward.json, summary.csv, and summary.json. No alternate pass field is emitted. The runtime applies this rule to the staged reward adapter after unsealing; encrypted archives and partial-credit rubrics remain unchanged. The updated runtime also removes native pass decisions from published verifier diagnostics. Jobs created by older runtimes cannot be resumed under the new policy: use a fresh job name. Old files are left untouched; summarizing an old job recomputes the official metric but does not rewrite its logs or rewards. The summary records metric_definition: score-gt-0.999. Historical results using per-task thresholds, == 1.0, or >= 0.999 must be recomputed from raw scores before comparison. See the scoring guide and runtime update.

Runtime rollout: the runtime update is pending. Before reporting results, check that your summary records metric_definition: score-gt-0.999; older runtimes use the legacy rule. Use the runtime summary for the fixed 97-task denominator; Harbor may aggregate only attempted trials.

Integrity

The root README.md is a mutable dataset card and is intentionally outside checksums.sha256. Task files (including task-level READMEs), manifests, registries, and verification tools remain checksummed. Payload changes require regenerating their checksum entries; editing this card does not.

Citation

bibtex
@misc{apodex11,
  title         = {Apodex 1.1: Scaling Agentic Intelligence for Complex Work},
  author        = {{Apodex Team}},
  year          = {2026},
  eprint        = {2608.23283},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2608.23283}
}

@misc{frontierchallenge,
  title         = {FrontierChallenge: Evaluating Scientific Workflow Completion},
  author        = {Liangcai Su and Zhaopeng Feng and Zhuo Chen and Zhen Zhang
                   and Xiang Lin and Ruilin Li and Handuo Zhang and Ning Wang
                   and Kailong Wen and Yueqi Guo and Feng Xing and Yiling Guo
                   and Chenxiong Qian and Simon Shaolei Du and Lidong Bing
                   and Xinyu Wang},
  year          = {2026},
  eprint        = {2608.24979},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2608.24979}
}