datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
varroa_mmdet_yolo_protocol_runsinteraction_protocol
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
This repository contains the processed experimental datasets used in:
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate. Pratik S. Sachdeva and Tom van Nuenen. COLM 2026.
Dataset contents
The experiments/ directory contains Parquet datasets used to reproduce the
figures and analyses in the paper. It includes:
synchronous head-to-head debates;
round-robin head-to-head debates;… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/interaction_protocol.honesty-index
The Kerne Honesty Index
What each synthetic dollar advertises, next to what it actually paid.
Advertised APY versus realized APY for 21 synthetic dollar vaults, recomputed hourly from
ERC-4626 share price growth on chain, and signed.
The realized column is not taken from anybody's dashboard. It is measured directly from the vault
contract: convertToAssets(10**decimals) read at two block heights, divided by 10**asset_decimals,
annualized over the real elapsed time between those… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/honesty-index.solana-yield-honesty
Solana Honesty Index
What each Solana stablecoin product says it pays, next to what it actually
paid, measured from a share price rather than from a claim.
Snapshot generated 2026-09-21T13:24:04.031Z. Window 30 days.
13 products across 3 protocols,
13 comparable, 0 published but not
comparable. Realized figures: 5 by issuer_share_price_history, 2 by onchain_share_price, 6 by issuer_share_price_observed.
product
advertised
realized
gap
delivered
realized method
Kamino… See the full description on the dataset page: https://huggingface.co/datasets/kerne-protocol/solana-yield-honesty.NOMOS-GEO-Audit-Protocol
NOMOS GEO Audit Protocol
A repeatable way to test what AI systems say about an organisation and whether the evidence supports it
GEO means Generative Engine Optimization. This six-language candidate protocol turns that discipline into an auditable process using the GEO-1000 method, canonical questions, truth packs, evidence requirements, scoring logic, correction steps and revalidation records.
Start reading: Open the English PDF · Choose one of six languages · Cite… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GEO-Audit-Protocol.agentic-publication-protocol-dataset
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 quantum-physics papers.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness, grounding, and honesty (1-10).
Evaluator: Codex CLI, gpt-5.5, reasoning effort xhigh… See the full description on the dataset page: https://huggingface.co/datasets/phynics/agentic-publication-protocol-dataset.bio-faiss-longevity-v1
bio-faiss-longevity-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-longevity-v1.neophyte-faiss-index-v1
neophyte-faiss-index-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
index.info.json: (optional) dimensions, index type, faiss version.
Build provenance
Chunking: hierarchical (section→paragraph→~480-token chunks, ~15% overlap)
Embedder:… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/neophyte-faiss-index-v1.real01c-insert-marker-d1-blind-dagger-r1-25k-protocolThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 112,
"total_frames": 32256,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:112"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01c-insert-marker-d1-blind-dagger-r1-25k-protocol.NOMOS-GBO-Audit-Protocol
NOMOS GBO Audit Protocol
Prove what an AI agent did, what authorised it, which evidence supports the judgement and whether it could be stopped.
Saying that an AI agent followed its instructions is not evidence. The NOMOS GBO Audit Protocol turns Generative Behavior Optimization (GBO) into a practical method for testing authority, tool use, evidence, delegation, stopping, recovery and human control.
Start reading: Open the English PDF · Choose one of six editions
Kaan… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GBO-Audit-Protocol.DeFi-Protocol-Data-on-Ethereum-2023-2024Decentralized finance (DeFi) has emerged as a significant sector within the blockchain and cryptocurrency space. DeFi protocols enable users to access financial services without traditional intermediaries, offering a wide range of applications such as lending, borrowing, trading, and yield farming. Understanding user behavior, protocol interactions, and market trends is crucial for analyzing the DeFi ecosystem's dynamics and identifying opportunities for innovation and growth.
About… See the full description on the dataset page: https://huggingface.co/datasets/mriusero/DeFi-Protocol-Data-on-Ethereum-2023-2024.protocols-with-stepsAll protocols from https://github.com/protocolsio/protocols in text form with steps as json list
protocol-bench
Protocol-Bench
15 published IEEE 802.11 and 3GPP procedures with ground-truth safety verdicts — and, where a
property fails, the shortest counterexample trace that proves it.
Most reasoning benchmarks accept an answer. This one asks for a proof: if a model says a protocol
is broken, it must supply a trace that starts at the initial state, moves only along real
transitions, and ends in a genuinely violating state. Traces are replayed mechanically. A
plausible-sounding trace that… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/protocol-bench.defillama-protocols-scraper-sample-data
DefiLlama Protocols Scraper
Scrape all 7,000+ DeFi protocols from DefiLlama in one run — TVL, 1h/1d/7d TVL change, market cap, category, chains and links. Filter by chain, category and TVL. Schedule it daily to track the entire DeFi landscape.
What the actor scrapes
🦙 DefiLlama Protocols Scraper — Scrape All DeFi Protocols & TVL Data Scrape all 7,000+ DeFi protocols from DefiLlama in a single run and export them to JSON, CSV or Excel. This DefiLlama scraper… See the full description on the dataset page: https://huggingface.co/datasets/logiover/defillama-protocols-scraper-sample-data.bio-faiss-d1ckgpt-v1
bio-faiss-d1ckgpt-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-d1ckgpt-v1.Discord-Dialogues-Preprocessed-Luna-Protocol
Discord-Dialogues-Preprocessed-Luna-Protocol is a preprocessed fork of mookiezi/Discord-Dialogues, adapted for fine-tuning Qwen2.5-family models as part of the Luna Protocol project.
This dataset contains anonymized Discord conversations for training and evaluating realistic conversational AI models in a ChatML-friendly format. It is derived directly from mookiezi/Discord-Dialogues with two targeted preprocessing steps applied (see below) — the underlying conversations, filtering pipeline… See the full description on the dataset page: https://huggingface.co/datasets/fox3000foxy/Discord-Dialogues-Preprocessed-Luna-Protocol.ai_data_defi_protocols
💎 Freemium Feed Notice: This public Hugging Face feed provides verified daily samples.Need complete raw historical archives, custom company contact points, or real-time REST webhook feeds?🌐 Upgrade to Institutional Full Feeds at alphaville.space or email lead strategist Sloane Valentine at alphaville-insights@agentmail.to.
📊 DeFi Protocol Intelligence & Alpha
Curated & Maintained Autonomously by Sloane Valentine @ Alphaville InsightsContact: alphaville-insights@agentmail.to… See the full description on the dataset page: https://huggingface.co/datasets/Alphaville-Insights/ai_data_defi_protocols.bio-faiss-microbiome-v1
bio-faiss-microbiome-v1
A FAISS index + metadata for scientific retrieval
Contents
index.faiss: FAISS index (cosine w/ inner product).
meta.jsonl: one JSON per chunk; fields include chunk_id, paper_id, title, section, subsection, paragraph_index, keywords, boost.
Build provenance
Chunking: hierarchical (section→paragraph→~380-token chunks, ~15% overlap)
Embedder: bio-protocol/scientific-retriever (mean-pooled, L2-normalized)
Similarity: cosine via inner… See the full description on the dataset page: https://huggingface.co/datasets/bio-protocol/bio-faiss-microbiome-v1.ProtocolEC
ProtocolEC
A protocol-derived benchmark for complete clinical-trial eligibility-criteria (EC) generation.
ProtocolEC pairs each of 4,302 completed Phase III trials (22 therapeutic areas) with (i) its
ClinicalTrials.gov registry metadata and registry EC, and (ii) a more complete EC set extracted
from the trial's protocol PDF. Protocol EC contain roughly twice the criteria and words of the
registry EC. Splits are 80/10/10, stratified by therapeutic area.
Split
Trials… See the full description on the dataset page: https://huggingface.co/datasets/Konghao/ProtocolEC.real01c-insert-marker-d1-blind-dagger-r2-baseline-uniform-ours-iql-protocolThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 116,
"total_frames": 36234,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:116"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01c-insert-marker-d1-blind-dagger-r2-baseline-uniform-ours-iql-protocol.omnimind-brain-organoids-four-protocols-baseline
OmniMind — Brain organoids four protocols (Naas et al. 2024)
Human brain organoid single-cell RNA-seq across four differentiation protocols.
Source
Zenodo: https://zenodo.org/records/15873911
GitHub: https://github.com/jn-goe/brain_organoids_four_protocols
Paper: Naas et al. 2024, bioRxiv, DOI: 10.1101/2024.11.15.623576
Files
brain_organoids_combined_cell_metadata.parquet — 69,794 cells, 24 metadata columns.… See the full description on the dataset page: https://huggingface.co/datasets/fabricioslv/omnimind-brain-organoids-four-protocols-baseline.eval_act_rac_protocolThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 27,
"total_frames": 19982,
"total_tasks": 1,
"chunks_size": 1000000,
"data_files_size_in_mb": 10000,
"video_files_size_in_mb": 50000,
"fps": 30,
"splits": {
"train": "0:27"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kd-forge/eval_act_rac_protocol.clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1Clinical Quad Investigator Turnover Training Reset Protocol Deviations Data Lag v0.1
Each row is a site monthly snapshot.
Core quad
Investigator turnoverTraining resetProtocol deviationsData lag
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-investigator-turnover-training-reset-protocol-deviations-data-lag-v0.1.clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1
Clinical Quad Enrollment–Protocol Deviations–Site Variance–Endpoint Integrity v0.1
What this is
A quad-coupling dataset for trial collapse driven by the interaction of:
enrollment pattern changes
rising protocol deviations
site-to-site variance
endpoint integrity degradation
Task
Input: one quad state rowOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters
Trials often fail through operational pressure:
recruitment becomes spiky or slow… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.1.Discord-Dialogues-Preprocessed-Luna-Protocol-role-lunaclinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2Clinical Quad Enrollment Protocol Deviation Site Variance Endpoint Integrity v0.2
What this dataset does
It tests whether a model can detect when endpoint integrity degrades under four coupled operational pressures.
Quad nodes
enrollment_pattern
protocol_deviation_rate
site_variance_level
endpoint_integrity
Labels
0 coherent
endpoints clean
enrollment stable
deviations not high
site variance not high
1 tradeoff
strain exists
endpoint softens or system drifts… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-enrollment-protocol-deviation-site-variance-endpoint-integrity-v0.2.square-d0-dagger-blind-v3u-shared-r5-protocolThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 363,
"total_frames": 67760,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:363"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/square-d0-dagger-blind-v3u-shared-r5-protocol.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2Clinical Quad Population Shift Protocol Deviation Site Variance Endpoint Fragility v0.2
What this dataset does
It tests whether a model can detect when clinical trial endpoints lose credibility under quad coupling.
Quad nodes
population_shift
protocol_deviation_rate
site_variance_level
endpoint_fragility
Labels
0 coherent
Stable population
Low deviations
Low site variance
Endpoint robust
1 tradeoff
Some drift exists
Endpoint still usable
Risk is present but not… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.2.clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1
Clinical Quad Population Shift × Protocol Deviation × Site Variance × Endpoint Fragility v0.1
What this is
A quad-coupling dataset for trial collapse that happens when:
the enrolled population drifts from the intended cohort
protocol deviations rise
site-to-site variance widens
the primary endpoint is fragile to measurement or baseline imbalance
Task
Input: one row describing the quad stateOutput: label
0 — Stable1 — Drift2 — Collapse
Why it matters… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-population-shift-protocol-deviation-site-variance-endpoint-fragility-v0.1.clinical-quad-site-training-protocol-complexity-error-rate-data-usability-v0.1Clinical Quad Site Training Protocol Complexity Error Rate Data Usability v0.1
Each row is a site week snapshot.
Core quad
Site training intensityProtocol complexityOperational error rateData usability
Target
label_data_collapse_next_60d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
