apodex
Datasets
All datasets matching “apodex”FrontierChallenge
FrontierChallenge
FrontierChallenge provides 97 scientific workflow tasks with plaintext
English instructions, inputs, Harbor definitions, domain labels, and the
redistributable open runtime image.
Path
Contents
manifest.jsonl
Dataset Viewer rows with task ID, taxonomy, difficulty, runtime, and instruction
tasks/<task-id>/
instruction.md, task metadata, environment definition, and agent-visible inputs
images/
Verified linux/amd64 Docker archive for the 81… See the full description on the dataset page: https://huggingface.co/datasets/apodex/FrontierChallenge.FrontierChallenge-reference
FrontierChallenge reference data
FrontierChallenge reference data provides 97 authenticated, encrypted
verifier archives.
Path
Contents
tasks/<task-id>/verifier.fcref
encrypted tests/: grader, rubric, fixtures, validation code, and reference outputs
manifest.jsonl
archive paths, sizes, and SHA-256 commitments
source_registry.json
release binding shared with GitHub and the solve dataset
tools/
integrity checker and standalone unsealer
The archive password is… See the full description on the dataset page: https://huggingface.co/datasets/apodex/FrontierChallenge-reference.Deep-Research-Benchmarks
Deep Research Benchmarks
Password-protected bundle of the public deep-research benchmarks used by AgentHarness to evaluate Apodex-1.0 in standard ReAct mode.
Download
wget https://huggingface.co/datasets/apodex/Deep-Research-Benchmarks/resolve/main/deep_research_benchmarks_260607.zip
unzip -P 'apodex*()_2026' deep_research_benchmarks_260607.zip
rm deep_research_benchmarks_260607.zip
Single quotes around the password are required — it contains *, (, ).
After… See the full description on the dataset page: https://huggingface.co/datasets/apodex/Deep-Research-Benchmarks.autobench-biomedical-verification
AutoBench Biomedical Verification
Frozen evaluation rows for AutoBench.
