datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
exodus-endpointsreasoning-duplication-endpoint-statex402-endpoint-readinessscvd.store x402 endpoint readiness corpus
scvd.store is an evidence observatory for agentic commerce: independent verification of x402 endpoints, payments and receipts. Before an agent pays an x402 endpoint, we check that it can be paid. After it pays, we check the signed receipt. Over time we watch endpoints and publish a dated, signed corpus. Sellers use it to prove a door works; buyers use it before spending. Every artifact is signed, expires, and names what we did not see. Not escrow, not… See the full description on the dataset page: https://huggingface.co/datasets/keeper-scvd/x402-endpoint-readiness.atlas-20-information-barrier-and-the-prompt-endpoint
ATLAS report 20: does the information barrier suppress verification, and where does the prompt line end?
Complete raw products of ATLAS rl-training report 20 (GitHub issue #43).
Two new selector surfaces over every question of the canonical LiveCodeBench
(175) and GPQA (198) validation sets, each question with all eight of its
cached candidates revealed. Both carry report 18's comparison closing and
change only the system message:
barrier relaxed — the finalization paragraph's… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-20-information-barrier-and-the-prompt-endpoint.hf-inference-endpoint-benchmarks
Raw benchmark result files
Raw JSON outputs from the sessions described in benchmarking-methodology.md. Model: Qwen3.5-4B family, hf-endpoints deployed via the configs documented in cli-and-api.md.
Short-prompt decode comparison (valid metric — prompt negligible vs output, see trap #1 in methodology doc)
File
Setup
llamacpp_results.json
llama.cpp, GGUF Q8_0, MTP, A10G
vllm_results.json
vLLM, FP8-dynamic, MTP, A10G
vllm_bf16_results.json
vLLM, bf16… See the full description on the dataset page: https://huggingface.co/datasets/LostGentoo/hf-inference-endpoint-benchmarks.clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.dataset-viber-image-generation-preference-inference-endpoints-battle-flux
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-image-generation-preference-inference-endpoints-battle-flux.mobile_demo_20260707_130921_edited_no_wheel_half_endpointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"x.vel",
"y.vel",
"theta.vel",
"endpoint_state"
],
"shape": [
4
]
},
"observation.state": {
"dtype": "float32"… See the full description on the dataset page: https://huggingface.co/datasets/rainbowrobotics/mobile_demo_20260707_130921_edited_no_wheel_half_endpoint.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1
Clinical Quad Enrollment–Protocol Deviations–Site Variance–Endpoint Integrity v0.1
What this is
A quad-coupling dataset for trial collapse driven by the interaction of:
enrollment pattern changes
rising protocol deviations
site-to-site variance
endpoint integrity degradation
Task
Input: one quad state rowOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters
Trials often fail through operational pressure:
recruitment becomes spiky or slow… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1.clinical-quad-adjudication-drift-endpoint-reclassification-timing-pressure-v0.1Clarus Clinical Quad Coupling Adjudication Drift Endpoint Reclassification Timing Pressure v0.1
What this dataset isThis dataset tests whether a model can detect endpoint adjudication drift driven by four interacting nodes.
Quad coupling nodes
Clustered endpoint reclassification
Source data delay or missing uploads
Exposure or dose documentation gaps
Governance or interim analysis pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-adjudication-drift-endpoint-reclassification-timing-pressure-v0.1.chembl-endpoint-context10-clean-labels
Clean ChEMBL Endpoint Context-10 Prediction Dataset
This dataset contains rendered question/answer examples from clean ChEMBL endpoint records. Each example gives 10 labeled reference molecules from one endpoint and asks for the clean label of one held-out query molecule from the same endpoint.
Sampling Procedure
The source is ChEMBL 36. Continuous labels are median pchembl_value per molecule-endpoint. Binary labels are exact whitelist text labels only; rows with pChEMBL… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/chembl-endpoint-context10-clean-labels.sequenced_endpoint_0clarus-endpoint-definition-integrity-v0.1.1Clarus Endpoint Definition Integrity v0.1.1
What this measures
You stop endpoint-label reuse from hiding definition changes
You force the model to name the missing endpoint details
Task per row
Pick A or B
Write the required note with the key definition difference
Gold support
required_keywords is a "|" list of key terms your note must include
Predictions file
CSV columns
sample_id
predicted_option
predicted_note
chembl-endpoint-pair-clean-labels
Clean ChEMBL Endpoint Pair Prediction Dataset
This dataset contains rendered question/answer examples from clean ChEMBL endpoint pair records.
Answers are sourced from clean aggregated labels: median pChEMBL for continuous endpoints and exact-whitelist 0/1 labels for binary endpoints.
Sampling Procedure
The source is ChEMBL 36. Continuous endpoint labels use non-null pchembl_value rows aggregated by median for each molecule-endpoint. Binary labels use exact whitelist text… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/chembl-endpoint-pair-clean-labels.dataset-viber-chat-generation-preference-inference-endpoints-battle
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/davidberenstein1957/dataset-viber-chat-generation-preference-inference-endpoints-battle.endpoint-framework-transition-integrity-v0.1Clarus Clinical Narrative Integrity v0.1.1
What this dataset tests
You track meaning as it moves across sections of a clinical paper
You detect when claims shift strength or scope without justification
You flag narrative drift introduced by section transitions
Scope
One document
Multiple internal sections
Abstract
Methods
Results
Discussion
Conclusion
One publication lifecycle
How this differs from v0.1
v0.1 tests internal consistency within a single narrative fragment
v0.1.1 tests… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/endpoint-framework-transition-integrity-v0.1.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2Clinical Quad Enrollment Protocol Deviation Site Variance Endpoint Integrity v0.2
What this dataset does
It tests whether a model can detect when endpoint integrity degrades under four coupled operational pressures.
Quad nodes
enrollment_pattern
protocol_deviation_rate
site_variance_level
endpoint_integrity
Labels
0 coherent
endpoints clean
enrollment stable
deviations not high
site variance not high
1 tradeoff
strain exists
endpoint softens or system drifts… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2.clarus-endpoint-framework-transition-integrity-v0.1Clarus Clinical Narrative Integrity v0.1
What this dataset tests
You track internal consistency within a single clinical narrative
You detect contradictions between claims, definitions, and reported results
You flag narrative overreach without changing formal frameworks
Scope
One document
One reporting layer
One definition space
Typical failure modes
Outcome language exceeds reported metrics
Safety claims conflict with tabulated data
Efficacy statements ignore stated limitations… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clarus-endpoint-framework-transition-integrity-v0.1.clinical-quad-endpoint-adjudication-timing-missingness-conmed-bias-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Timing Missingness Conmed Bias v0.1
What this dataset isThis dataset tests whether a model can detect endpoint adjudication risk driven by four interacting nodes.
Quad coupling nodes
Measurement or assessment timing drift
Concomitant medication or treatment timing
Data missingness affecting classification
Governance constraints such as committee deadlines or interim reviews
Input
One vignette in prompt
OutputReturn strict JSON… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-timing-missingness-conmed-bias-v0.1.mobile_demo_20260707_130921_edited_no_wheel_no_odom_half_endpointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"x.vel",
"y.vel",
"theta.vel",
"endpoint_state"
],
"shape": [
4
]
},
"observation.state": {
"dtype": "float32"… See the full description on the dataset page: https://huggingface.co/datasets/rainbowrobotics/mobile_demo_20260707_130921_edited_no_wheel_no_odom_half_endpoint.endpointDetectDatasetUnslotclinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2Clinical Quad Population Shift Protocol Deviation Site Variance Endpoint Fragility v0.2
What this dataset does
It tests whether a model can detect when clinical trial endpoints lose credibility under quad coupling.
Quad nodes
population_shift
protocol_deviation_rate
site_variance_level
endpoint_fragility
Labels
0 coherent
Stable population
Low deviations
Low site variance
Endpoint robust
1 tradeoff
Some drift exists
Endpoint still usable
Risk is present but not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1
Clinical Quad Population Shift × Protocol Deviation × Site Variance × Endpoint Fragility v0.1
What this is
A quad-coupling dataset for trial collapse that happens when:
the enrolled population drifts from the intended cohort
protocol deviations rise
site-to-site variance widens
the primary endpoint is fragile to measurement or baseline imbalance
Task
Input: one row describing the quad stateOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1.clinical-quad-endpoint-adjudication-bias-missingness-timing-pressure-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Bias Missingness Timing Pressure v0.1
What this dataset isThis dataset tests whether a model can detect endpoint adjudication bias created by four interacting forces.
Quad coupling nodes
Endpoint rate or classification shift
Data latency or packet incompleteness
Concomitant exposure or contextual missingness
Governance pressure such as interim look, submission, earnings, or regulator briefing
Input
One vignette
OutputReturn… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-bias-missingness-timing-pressure-v0.1.clinical-quad-safety-endpoint-sponsor-subgroup-collapse-v0.1
Clinical Quad Safety–Endpoint–Sponsor–Subgroup Collapse v0.1
What this is
A quad-coupling dataset that models collapse when four forces lock together:
Safety signal strength
Endpoint outcome state
Sponsor pressure intensity
Subgroup fragility
Task
Input: a quad state rowOutput: stability label
Labels
0 — Stable1 — Drift2 — Collapse
Core idea
Trials often do not “fail cleanly”.
A weak-to-moderate safety signal plus a missed endpoint can trigger… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-endpoint-sponsor-subgroup-collapse-v0.1.clinical-quad-evidence-drift-endpoint-signal-claim-language-certainty-narrative-break-v0.1What this repo does
This dataset models narrative continuity break in clinical trial summaries. It predicts when the interaction between evidence consistency, endpoint signal strength, claim strength, and certainty language indicates that the written narrative has drifted away from the underlying trial results.
Core quad
evidence_consistency_index
endpoint_signal_strength_index
claim_strength_index
certainty_language_index
Prediction target
label_narrative_break
Row structure
Each row… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-evidence-drift-endpoint-signal-claim-language-certainty-narrative-break-v0.1.endpoints-userstoriesThis dataset contains a list of endpoints with parameters and a column for the user story of that endpoint
inference-endpoints-structured-generation-multiple
Dataset Card for inference-endpoints-structured-generation-multiple
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/inference-endpoints-structured-generation-multiple/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/inference-endpoints-structured-generation-multiple.endpoint_datasetclinical-quad-recruitment-inclusion-severity-endpoint-dilution-v0.1Clinical Quad Recruitment Inclusion Severity Endpoint Dilution v0.1
Each row is a site week snapshot.
Core quad
Recruitment speedInclusion criteria tightnessPopulation severityEndpoint dilution
Target
label_primary_miss_next_60d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
