datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
appworld-qwen35-4b-total-237-audited-jh-epoch2
appworld-qwen35-4b-total-237-audited-jh-epoch2
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.3640625
Action score: 0.4328125
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch6
appworld-qwen35-4b-total-237-audited-jh-epoch6
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.37578125
Action score: 0.421875
Valid samples: 320/320
appworld-qwen35-4b-total-237-audited-jh-epoch8
appworld-qwen35-4b-total-237-audited-jh-epoch8
Portable process-evaluation output. metadata.json is the lightweight source
for aggregate results; the JSONL files are directly loadable; and
artifacts.tar.gz losslessly preserves the original run directory.
Reasoning score: 0.384375
Action score: 0.4390625
Valid samples: 320/320
slither-audited-smart-contractsThis dataset contains source code and deployed bytecode for Solidity Smart Contracts that have been verified on Etherscan.io, along with a classification of their vulnerabilities according to the Slither static analysis framework.FrontierOR-Audited-92
FrontierOR Audited 92
This is a derived, evaluator-oriented release of 92 cases from
SmartOR/FrontierOR, pinned to
upstream dataset commit 37ccd8b6dca3bf7f4e0c58941a6ed156832a6d9e.
The original 180-case release was audited for coherence between the problem
description/formulation, reference Gurobi implementation, solution schema, bundled
solutions, and feasibility checker. This release contains the 92 cases that were not
assigned BLOCK_RELEASE; 91 checkers were hardened and the… See the full description on the dataset page: https://huggingface.co/datasets/LeoJiangOR/FrontierOR-Audited-92.FrontierOR-Audited-180
FrontierOR Audited 180
This is an evaluator-oriented derivative of
SmartOR/FrontierOR, pinned to
upstream commit 37ccd8b6dca3bf7f4e0c58941a6ed156832a6d9e. The released descriptions, formulations, Gurobi
implementations, solution schemas, reference solutions, and feasibility checkers were
audited and repaired as one evaluation contract.
Final status
Gate
Result
Evaluator-ready directory IDs
180/180
Independent canonical cases
179
Transparent… See the full description on the dataset page: https://huggingface.co/datasets/LeoJiangOR/FrontierOR-Audited-180.pai-interview-a2-human-audited-quarantine-vizThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 12,
"total_frames": 5773,
"total_tasks": 12,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:12"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/geonmin-kim/pai-interview-a2-human-audited-quarantine-viz.suds-full-routine-92-camera-audited-v2fa-ds-b1k-task01-filtered-recovery-audited-4739b3054a80-caecf22e676d
ds-b1k-task01-filtered-recovery-audited-4739b3054a80
Private fleet archive. Canonical recipe: evaluations/2026-09-22_b1k_task01_three_source_distribution_eai
Tier: historical audit input reference; consumed subset fingerprinted
Use the exact recorded revision and verify SHA256SUMS. This package is a snapshot, not a live directory mirror.
Slither-Audited-Solidity-QA
Dataset Card for "Simple-Solidity-Slither-Vulnerabilities"
More Information needed
ChartVerse-RL-10K-auditedgsm-noop-audited
GSM-NoOp (audited)
NoOp distractor clauses for GSM-Symbolic questions, audited for genuine irrelevance by an independent model (GPT-5.5), plus per-model evaluation results.
Accompanies the post Revisiting GSM-Symbolic: Do 2026 Frontier Models Still Fail at Confounded Grade School Math? and the code at github.com/BenSturgeon/gsm-symbolic-revisited.
Why
The GSM-Symbolic paper (Mirzadeh et al., Apple, ICLR 2025) reported that adding an irrelevant "NoOp" clause collapses… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/gsm-noop-audited.fa-ds-b1k-task01-v14-full-task-audited-e05e32b95a5f-78bfe49b6bbc
ds-b1k-task01-v14-full-task-audited-e05e32b95a5f
Private fleet archive. Canonical recipe: evaluations/2026-09-22_b1k_task01_three_source_distribution_eai
Tier: historical audit input reference; consumed subset fingerprinted
Use the exact recorded revision and verify SHA256SUMS. This package is a snapshot, not a live directory mirror.
fa-ds-b1k-task01-human-audited-f342472e104c-31509c34abba
ds-b1k-task01-human-audited-f342472e104c
Private fleet archive. Canonical recipe: evaluations/2026-09-22_b1k_task01_three_source_distribution_eai
Tier: historical audit input reference; consumed subset fingerprinted
Use the exact recorded revision and verify SHA256SUMS. This package is a snapshot, not a live directory mirror.
fa-ds-b1k-task01-full-recovery-success-audited-f5cb94eb9a0-36a747d8293c
ds-b1k-task01-full-recovery-success-audited-f5cb94eb9a00
Private fleet archive. Canonical recipe: evaluations/2026-09-22_b1k_task01_three_source_distribution_eai
Tier: historical audit input reference; consumed subset fingerprinted
Use the exact recorded revision and verify SHA256SUMS. This package is a snapshot, not a live directory mirror.
Enhanced-Slither-Audited-Solidity-QA
Dataset Card for "Enhanced-Slither-Audited-Solidity-QA"
More Information needed
TruthfulQA-Audited
TruthfulQA-Audited
Datasets accompanying an anonymous NeurIPS 2026 Evaluations & Datasets
Track submission on surface-form leakage in binary-choice truth
benchmarks. The release contains three related artifacts:
TruthfulQA-476
Cleaned subset of binary-choice TruthfulQA, with surface-form leakage
removed via an audit-and-prune procedure.
canonical_label: TruthfulQA-476
theta: 0.53
n_pairs: 476
audit AUC: 0.528
derived from: binary-choice TruthfulQA (790 pairs)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymNeurIPS2026submission/TruthfulQA-Audited.indian-scam-sms-synthetic-audited
Indian Scam SMS (synthetic, audited)
1,580 short messages that imitate SMS and WhatsApp scams and their genuine look-alikes in Indian
English, Hindi (Devanagari), Hinglish and four Roman-script code-mixed styles (Tamil, Telugu, Bengali,
Marathi with English). Every row was written by a large language model and then audited for label noise.
It exists to train and stress-test scam detectors on the hard negatives that public datasets lack:
real-looking bank, courier, bill and job… See the full description on the dataset page: https://huggingface.co/datasets/Ridham115/indian-scam-sms-synthetic-audited.
