datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/harborframework/parity-experiments.p2-etf-risk-parity-resultsdabstep-parity-results
DABstep Parity Experiment
Overview
Adapter: DABstep (Data Agent Benchmark for Multi-step Reasoning)
Agent: claude-code
Model: anthropic/claude-haiku-4-5
Tasks: 130 (stratified sample from 450 default-split tasks: 21 easy + 109 hard, seed=42)
Trials: 4
Agent Timeout: 1800 seconds
Oracle Results
460/460 tasks passed (100%).
Parity Results
Split
Original (fork)
Harbor
Delta
Overall
37.69 ± 0.89%
36.92 ± 0.63%
-0.77 pp
Easy (21)
84.52 ±… See the full description on the dataset page: https://huggingface.co/datasets/Hudx111/dabstep-parity-results.parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, interpretation, notes, etc.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├── original_parity/
├── harbor_parity/
├── oracle/
└── results_collection/… See the full description on the dataset page: https://huggingface.co/datasets/DarrenDong/parity-experiments.cryocare-v0.3.0-parity-evidence
cryoCARE v0.3 CPU parity evidence
This dataset contains only JSON and Safetensors. It preserves deterministic inputs,
the exact extracted Keras-layout tensor inventory, vendor TensorFlow outputs, native
PyTorch outputs, and numerical metrics. The original HDF5 is not uploaded; its exact
hash and size remain in vendor/vendor-capture.json and the model conversion record.
The corresponding native Safetensors package is
scitomo/cryocare-v0.3.0-synthetic-parity.
marin-qwen36-nemosci-terminus2-b200-parity-configsmsr_zhen_translation_parityTranslator Human Parity Data
Human evaluation results and translation output for the Translator Human Parity Data release,
as described in https://blogs.microsoft.com/ai/machine-translation-news-test-set-human-parity/.
The Translator Human Parity Data release contains all human evaluation results and translations
related to our paper "Achieving Human Parity on Automatic Chinese to English News Translation",
published on March 14, 2018.bfcl-paritybfcl_parity_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_20260604_192648large-model-agentic-eval-parity-tb21-aa
Large Model Agentic Eval Parity with AAII on TB2.1
This repository contains the launch configurations, Harbor results, trajectories, and analysis artifacts for marin-community/marin#8261.
The campaign evaluates Qwen3.5-122B-A10B-FP8, DeepSeek-V4-Flash-0731, Nemotron 3 Ultra 550B A55B NVFP4, and GLM-5.2 AWQ INT4 against Artificial Analysis Terminal-Bench 2.1 results.
Redaction
Resolved Harbor records and captured terminal output contained signed endpoint URLs and… See the full description on the dataset page: https://huggingface.co/datasets/penfever/large-model-agentic-eval-parity-tb21-aa.DCAgent2_bfcl-parity_laion_GLM-4_7-inferredbugs-sandboxes-maxeps-131k_20260226_044018bfcl_parity_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_214919Harbor-Parity-Test-ARC-AGI-2bfcl_parity_GLM_4_7_swesmith_sandboxes_with_tests_oracle_verified_120s_maxeps_1015c97b3marin-qwen36-nemosci-terminus2-b200-parity-artifacts
Marin Qwen3.6 and Nemosci Terminus-2 parity artifacts
This repository contains the deduplicated Harbor outputs, trajectories, operational evidence, and analysis for three Terminus-2 parity evaluations of Qwen3.6-35B-A3B and Nemosci Qwen3-32B on CoreWeave GB200 GPUs.
Each model and benchmark is packaged separately:
Archive
SHA-256
qwen3.6-dev-set-v2.tar.zst
884f78227e9838b9e762a46ec769141a1369b4b036792e6f196c07f8ab9ac465
qwen3.6-swebench.tar.zst… See the full description on the dataset page: https://huggingface.co/datasets/laion/marin-qwen36-nemosci-terminus2-b200-parity-artifacts.parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, directory structure, trajectory interpretation, notes, etc. This should be DIFFERENT than the adapter REAMDE.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├──… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/parity-experiments.wrapped-asset-parity
Wrapped-Asset Parity Ledger
Council of AI · CSOAI Ltd (UK #16939677) · live source https://councilof.ai/interop/wrapped-asset-parity-latest.json · door GET https://councilof.ai/api/wrapper?id=<pair> (free &preview=1, x402 402 challenge for the signed card)
One row per bridged or custodial wrapper pair, read from public RPC with no key at provider-reported finalized blocks: the wrapped token's totalSupply() on its chain and, where an escrow exists, the canonical token's… See the full description on the dataset page: https://huggingface.co/datasets/csoai/wrapped-asset-parity.DCAgent2_bfcl-parity_mlfoundations-dev_codeforces-sandboxes-traces-terminus-2_2e80c2215DCAgent2_bfcl-parity_laion_open-thoughts-4-code-qwen3-32b-annotated_20260227_205932DCAgent2_bfcl-parity_laion_claude-4-5-sonnet-thinking-stackexchange-overflow-32c7ad75bcControlNet-Shadows
Dataset Card for "Shadow-Dataset-ControlNet"
More Information needed
DCAgent2_bfcl-parity_DCAgent_freelancer-random-instruction-filter-traces-termin32fb3559quant-vs-api-parity-harness
🔬 quant-vs-api-parity-harness
Is your local quant actually as good as the full-precision API? This toolkit answered
that question for a 284B MoE compressed to 2.9 bpw — verdict: indistinguishable on the
real serving path (90.8% token-identical, 240/240 paired-QA parity, deep-derivation
parity). 🎯 But the real product is the method: it caught 9 bugs in our own
instruments before they could lie to us — including one that had us convinced the quant
had lost factual recall when… See the full description on the dataset page: https://huggingface.co/datasets/Kevletesteur/quant-vs-api-parity-harness.parity-experiments
Upload your Adapter Oracle and Parity results
This dataset saves the oracle and parity experiment logs for adapters. Please upload them according to the following format and draft a PR.
adapters/
└── {adapter_name}/
├── README.md # Results overview, interpretation, notes, etc.
├── config.yaml # The yaml file that can be directly used to run parity experiments in Harbor.
├── original_parity/
├── harbor_parity/
├── oracle/
└── results_collection/… See the full description on the dataset page: https://huggingface.co/datasets/Ziruo03/parity-experiments.fine-tune-vs-rag-parity-index
Fine-Tune vs RAG — parity retrieval index
The exact FAISS index the
Fine-Tune vs RAG benchmark
used for its rag-parity arm, published so the
live demo
retrieves over the same passages the report did.
Not for clinical use. Not medical advice. Not a medical device. This is
exam-explanation text from a public benchmark dataset, chunked for retrieval
research. It contains the textual noise and errors documented in the report.
What is in it
File
Contents… See the full description on the dataset page: https://huggingface.co/datasets/vireshk/fine-tune-vs-rag-parity-index.topaz-v0.2.4-parity-evidence
Topaz v0.2.4 tiled CUDA parity evidence
This dataset preserves the bounded historical tiled-inference realization used
by Scitomo's Topaz parity comparison. It contains a deterministic nonconstant
float32 MRC fixture, Topaz 0.2.4 CUDA outputs for vendor aliases unet-3d-10a
and unet-3d-20a, launcher logs, and immutable provenance.
The capture ran the full Topaz CLI from the clean official source checkout at
c6dde54398875dcc6a210f83de019c9165fb474c, source distribution SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/scitomo/topaz-v0.2.4-parity-evidence.DCAgent2_bfcl-parity_laion_rl_base-code-contests-900s-160_20260227_205921DCAgent2_bfcl-parity_DCAgent2_test2-tbench-dev-71-qwen3-8b-8nodes-sync_20260227_045802DCAgent2_bfcl-parity_DCAgent2_test2-tbench-dev-71-qwen3-8b-8nodes-sync_20260227_032036osworld-native-parity-runs
OSWorld Native Parity Runs
This dataset stores large native OSWorld parity run archives that are too large for the shared Harbor parity-experiments dataset.
Archives
fulltask_20260621-172216/attempt_1/osworld-native-fulltask_20260621-172216-attempt1-349of361.tar.zst
Source run: /home/servermacadmin/osworld-parity/parity_results/fulltask_20260621-172216/attempt_1
Source upstream: xlang-ai/OSWorld at fe8c78e
Tasks: OSWorld-Verified no-Google-Drive split, 361 task… See the full description on the dataset page: https://huggingface.co/datasets/josancamon/osworld-native-parity-runs.
