apodex/FrontierChallenge
FrontierChallenge FrontierChallenge provides 97 scientific workflow tasks with plaintext English instructions, inputs, Harbor definitions, domain labels, and the redistributable open runtime image. Path Contents manifest.jsonl Dataset Viewer rows with task ID, taxonomy, difficulty, runtime, and instruction tasks/<task-id>/ instruction.md, task metadata, environment definition, and agent-visible inputs images/ Verified linux/amd64 Docker archive for the 81… See the full description on the dataset page: https://huggingface.co/datasets/apodex/FrontierChallenge.
FrontierChallenge
FrontierChallenge provides 97 scientific workflow tasks with plaintext English instructions, inputs, Harbor definitions, domain labels, and the redistributable open runtime image.
This repository contains no graders, rubrics, fixtures, or reference outputs. Those are stored as encrypted archives in the separate `apodex/FrontierChallenge-reference` dataset.
Use the FrontierChallenge runtime to download both datasets, verify their shared registry, load the image archive, run Harbor, and score a task. The evaluator host needs Python 3.12+ for the pinned Harbor 0.20.0; task containers keep their own frozen Python versions.
git clone https://github.com/ApodexAI/FrontierAgent.git
cd FrontierAgent/benchmarks/frontierchallenge
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install -e .
HF_TOKEN=hf_... ./scripts/setup.sh --track open
cp .env.example .env
# Fill model and judge credentials in .env before running.
./scripts/run_eval.sh \
--agent claude-code --model <model> \
--include-task-name task_011_cell_migration_wound_healingThe 16 task definitions that execute ORCA and their inputs are included normally. ORCA itself and an ORCA-configured image are never distributed; evaluators obtain ORCA from its official provider and follow the runtime repository's local-image tutorial.
Track membership follows what a task executes, not software named in supplied files. For example, task_098_orca_claisen_thermochemistry reads precomputed ORCA output without running ORCA, so it remains in the open track.
Verify a downloaded solve package with:
python tools/verify_dataset.pyMetrics
Official Pass Rate uses a strict score threshold: a task passes only when evaluation_complete == 1 and valid task_score > 0.999. Native per-task thresholds do not determine this metric. Use unrounded scores: exactly 0.999 does not pass; 0.9991 and 1.0 pass.
- Pass Rate (%) = 100 × number of completed tasks with valid
task_score > 0.999/ 97. - Score (0–100) = 100 × sum of valid, completed
task_scorevalues / 97. - Partial credit remains the native rubric score normalized to [0, 1]. Missing tasks, invalid scores, and incomplete evaluations contribute zero to the fixed denominator and must be reported separately; incomplete runs are not final benchmark results.
- Use one predeclared attempt per task. A subset may use its predeclared task count as denominator, but must be labeled as a subset, not the 97-task result.
There is only one pass field, passed, using the same strict threshold in reward.json, summary.csv, and summary.json. No alternate pass field is emitted. The runtime applies this rule to the staged reward adapter after unsealing; encrypted archives and partial-credit rubrics remain unchanged. The updated runtime also removes native pass decisions from published verifier diagnostics. Jobs created by older runtimes cannot be resumed under the new policy: use a fresh job name. Old files are left untouched; summarizing an old job recomputes the official metric but does not rewrite its logs or rewards. The summary records metric_definition: score-gt-0.999. Historical results using per-task thresholds, == 1.0, or >= 0.999 must be recomputed from raw scores before comparison. See the scoring guide and runtime update.
Runtime rollout: the runtime update is pending. Before reporting results, check that your summary records metric_definition: score-gt-0.999; older runtimes use the legacy rule. Use the runtime summary for the fixed 97-task denominator; Harbor may aggregate only attempted trials.
Integrity
The root README.md is a mutable dataset card and is intentionally outside checksums.sha256. Task files (including task-level READMEs), manifests, registries, and verification tools remain checksummed. Payload changes require regenerating their checksum entries; editing this card does not.
Citation
@misc{apodex11,
title = {Apodex 1.1: Scaling Agentic Intelligence for Complex Work},
author = {{Apodex Team}},
year = {2026},
eprint = {2608.23283},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.23283}
}
@misc{frontierchallenge,
title = {FrontierChallenge: Evaluating Scientific Workflow Completion},
author = {Liangcai Su and Zhaopeng Feng and Zhuo Chen and Zhen Zhang
and Xiang Lin and Ruilin Li and Handuo Zhang and Ning Wang
and Kailong Wen and Yueqi Guo and Feng Xing and Yiling Guo
and Chenxiong Qian and Simon Shaolei Du and Lidong Bing
and Xinyu Wang},
year = {2026},
eprint = {2608.24979},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2608.24979}
}