datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
assay-corpus
Assay corpus — audited RL environments with planted-defect ground truth
28 environments across 5 ecosystems, each audited by
Assay, an agentic auditor for RL environments and eval suites.
For the 17 environments whose ecosystem this repo owns, the corpus
ships the planted-defect ground truth and the full Environment Card. For the
11 it classifies as someone else's, it ships the verdict and
which probes could not run and why — and nothing else. See
What is not here.
The source… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/assay-corpus.assay-cupel-pose-error-v2-r7-mandatory-sandbox-rlassay-cupel-pose-error-v2-r6-sandbox-harness-rlweb-research-trajectories
Web-Research Agent Trajectories
The first open dataset from Assayo — an open rubric and method for judging the quality
of AI agent trajectories. (The name is from assay*: to test the purity of a metal.)*
An open rubric and a small, hand-built gold set for judging multi-step web-research
agent trajectories. A trajectory is the full record of an agent solving one task by
searching the web, reading sources, and answering with citations — the
think → act → observe → repeat →… See the full description on the dataset page: https://huggingface.co/datasets/Assayo/web-research-trajectories.assay-receipts
Assay Receipt Corpus
Signed internal receipts from an inference provider that is sometimes cheating, together
with the verdict an auditor reached on each one and the ground truth of which model actually
served the request.
Each row is a real receipt, not a summary statistic: it carries the prompt and output token
ids, the JL-projected sketch of the provider's hidden_states, the sign/rank invariants, and
an HMAC signature. With the gpt2 weights you can recompute the sketch… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/assay-receipts.assay-synthetic
assay-synthetic
Models trained on this: the Assay collection
-- six decision models from 149M to 27B. Berk/assay-4b
is the usual choice; Berk/assay-compiled-base
runs on a CPU.
Synthetic decision tasks generated by the Assay project.
Each item is a state (the facts), a typed question (bool, choice or score with
described options) and a label computed in code, so the labels are exact rather than
annotated. They were written to fix two measured weaknesses of decision models:… See the full description on the dataset page: https://huggingface.co/datasets/Berk/assay-synthetic.Assay-aware-BindingDB
Assay-aware BindingDB
Assay-aware BindingDB is a collection of protein–ligand binding
records organized by experimental assay type. Each row represents a BindingDB
reactant set and includes its measured affinity, source publication, original
experimental context, and an assay-specific structured description.
The complete dataset remains available as the full split. Four assay
configurations provide direct access to ITC, SPR, FPA, or RBA records, and 40
training-compatible… See the full description on the dataset page: https://huggingface.co/datasets/anonymousapple/Assay-aware-BindingDB.assay-qi
ASSAY-QI v2.0 — Quantum-Augmented BFSI Attack Suite
ASSAY-QI (Adversarial Safety Suite for AI — Quantum-Inspired) is a 1,273-prompt adversarial corpus for BFSI AI safety evaluation, generated using quantum circuit Born machine (QCBM) sampling and simulated annealing to target the decision boundary of BFSI safety classifiers.
Published by Zytra · Part of FinProof v1 · License: CC BY 4.0
What makes ASSAY-QI different
Standard adversarial datasets use template… See the full description on the dataset page: https://huggingface.co/datasets/Zytra/assay-qi.chembl-assay-transfer-k1assay-pose-replay-offset-tool-fewshot-rotating-expanded-v1-rlassay-pose-replay-offset-tool-fixed-demo-v1-rlassay-pose-replay-offset-tool-native-v1-rlassay-pose-replay-offset-tool-fewshot-rotating-v1-rlassay-pose-replay-offset-tool-fewshot-rotating-hardened-v1-rlassay-cupel-pose-error-v2-r8-all-frame-diagnostics-rlassaybenchassay-pose-replay-harness-v1-rlassay-cupel-pose-error-v2-r9-lowest-dispersion-rlassay-pose-replay-offset-tool-v1-rlchembl-assay-transfer-k1-cotchembl-assay-transfer-k10chembl-assay-transfer-k10-cotassay-cupel-pose-error-v2-r5-frame-harness-rlassay-band-v3-rlassay-cupel-pose-error-v1-rlassay-cupel-pose-error-v2-r1-rlassay-cupel-pose-error-v2-r3-rlassay-cupel-pose-error-v2-r4-harness-rlassay-cupel-pose-error-v1-r2-rlassay-cupel-pose-error-v2-r2-rl
