datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dim-discovery-archive
Geometry of Decision Making in Language Models
Abhinav Joshi · Divyanshu Bhatt · Ashutosh ModiNeurIPS 2025
This repository contains the official implementation/release for the NeurIPS 2025 paper Geometry of Decision Making in Language Models.
We study the internal decision-making processes of large language models through the lens of intrinsic dimension (ID), analyzing how hidden representations evolve across layers in a multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/Exploration-Lab/dim-discovery-archive.equity-perp-price-discovery
Equity and pre-IPO perpetual prices
Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets.
Contents
Table
Record
perpetual_mark_and_index_prices
An instrument's mark price, index price and basis at an observation time
Using the data
market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.discoverybenchData-driven Discovery Benchmark from the paper:
"DiscoveryBench: Towards Data-Driven Discovery with Large Language Models"
🔭 Overview
DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.discovery
Dataset Card for Discovery
Dataset Summary
Discourse marker prediction with 174 markers
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English
Dataset Structure
input : sentence1, sentence2,
label: marker originally between sentence1 and sentence2
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
Train/Val/Test
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sileod/discovery.agentic-drug-discovery-system
Agentic Drug Discovery System
This card describes the public 0.3.0.dev3 Agentic Drug Discovery System mirror.
Scope. The proposed eight-stage, long-horizon agentic drug discovery system remains a research scaffold rather than a completed public platform. Seven of eight planned atlases have no standalone public data, and the demonstrated continuous multi-stage program currently covers one disease/target slice traversed retrospectively.
It contains the executable control plane… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/agentic-drug-discovery-system.dblp-discovery-dataset
Dataset Card for DBLP Discovery Dataset (D3)
Dataset Summary
DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.discovery_discovery_promptsourcecircuit-discovery
Circuit Shotting Artifacts
This dataset stores large/generated artifacts externalized from Occupying-Mars/circuit-shotting.
The Git repo keeps source code, configs, lightweight reports, and current narrative documents. This HF dataset keeps generated datasets, raw result trees, pod artifact bundles, masks, dashboards, logs, adapters, and reference files that do not belong in normal Git history.
Artifact and Backup Locations
Issue 43 XCOMET Destructive… See the full description on the dataset page: https://huggingface.co/datasets/TokenBender/circuit-discovery.2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment.
field
value
experiment
LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.hf-coding-tools-traces-discovery
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
31 sessions, one per (tool, model, effort, thinking) configuration
9,022 query → response turns total (≈18,044 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.parkinsons-evidence-to-discovery-prioritisation
Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset
This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation.
Dataset Summary
The dataset integrates:
evidence-priority scores for PD prevention and disease-modification candidates;
pathway-to-intervention framework;
individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.csoai-layer0-discovery-cuts
CSOAI Layer 0 public discovery cuts
This repository preserves small, content-addressed observations of publicly served Council of AI discovery surfaces and a local gate map. It is a reproducibility aid, not a certification, training dataset, or record of revenue.
The first cut is e9c4df4e9383af62c5970cc3680cfcb5e3f095b93e8ad5824c807e74495fe145.json. Its filename is the SHA-256 of its exact bytes. It records 13 public MCP tool names, 21 x402 resource IDs, a public discovery post… See the full description on the dataset page: https://huggingface.co/datasets/Nicholastempleman/csoai-layer0-discovery-cuts.Reverse-circuit-discoverypd-discovery-benchmark-dashboard
Parkinson's Disease Discovery Benchmark Dashboard
Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery.
This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.2026-07-29-msm-philosophy-spec-focused-discovery
Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript.
Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c
Brief finding
No seed replicated. Ten seed archetypes were each run for three epochs. Under
the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.discoverybench
DiscoveryBench - Alias
A reformatted version of the original DiscoveryBench dataset for easier usage.
🤗 Original Dataset on HF
💻 GitHub Repository
📄 Paper (arXiv)
📁 Dataset Structure
The dataset consists of real and synthetic subsets:
Real Splits:
real_train
real_test
Synthetic Splits:
synth_train
synth_dev
synth_test
Each split contains a list of tasks with references to associated CSV datasets needed to answer the query. LLMs are expected to use the… See the full description on the dataset page: https://huggingface.co/datasets/nhop/discoverybench.scaling_law_discovery_results
Scaling Law Discovery Results Dataset
Results dataset for the paper: "Can Language Models Discover Scaling Laws?"
This dataset contains the complete collection of results from the Scaling Law Discovery (SLDBench) benchmark, where various AI agents attempt to discover mathematical scaling laws from experimental LLM training data.
🔗 Quick Links
Resource
Link
📄 Paper
arXiv:2507.21184
📊 Original Benchmark
SLDBench Dataset
🧪 Benchmark Code… See the full description on the dataset page: https://huggingface.co/datasets/pkuHaowei/scaling_law_discovery_results.red-pill-drug-discovery-formulation
🔴 RED-PILL
Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language
The first open instruction-tuning dataset for drug discovery & formulation development.
Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions.
⚡ Quick Start
from datasets import load_dataset
# Load the full dataset
ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.Original-circuit-discoveryzhangxian_1928_portrait_discovery_storm_society_modern_art_pioneer🖼️ Zhang Xian (1898–1936): The Lost 1928 Portrait and His Legacy in China's Modern Art Movement
Dataset title: zhangxian_1928_portrait_discovery_storm_society_modern_art_pioneer
Compiled by: HaruthaiAI (Hugging Face)
📄 Metadata JSON
Structured metadata file: zhangxian_portrait_metadata.jsonIncludes multi-layered context: discovery, attribution, movement, style, and linked documents.
📘 Overview (English)
This dataset documents a significant rediscovery: a 1928 portrait of Zhang… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/zhangxian_1928_portrait_discovery_storm_society_modern_art_pioneer.Science-Discoverydaily-paper-2026-08-03-deferred-tool-schema-discovery-cost
Cold-Start Tool Blindness: Measuring the Unknown-Unknown Cost of Deferred MCP Tool-Schema Loading in LLM Agent Harnesses
TL;DR — Deferred MCP tool-schema loading saves up to 91.4% of input tokens but silently fails to retrieve tools whose names and descriptions share no vocabulary with the agent's task description, recovering only 2 of 26 tools (7.7%) under unaligned queries — a failure mode we call cold-start tool blindness.
ThakiCloud AI Research · 2026-08-03 · 📝 Tech blog… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-08-03-deferred-tool-schema-discovery-cost.printable-coloring-product-discovery
PrintableFunnyPages — Printable Coloring Product Discovery & Buyer Intent Corpus
Dataset Summary
This dataset is an English-language product discovery and buyer-intent corpus created for PrintableFunnyPages, a digital printable shop offering downloadable coloring and activity resources.
It is designed to support research and experimentation in:
product discovery
semantic product retrieval
buyer-intent understanding
ecommerce search
recommendation and matching… See the full description on the dataset page: https://huggingface.co/datasets/PrintableFunnyPages/printable-coloring-product-discovery.discovery-2025
Discovery 2025 Dataset
Training data for a hyperlocal AI assistant for Wagholi, Pune (India).
Dataset Description
This dataset contains conversations in Hinglish (Hindi-English mix), Marathi, and English
for training a local discovery assistant that helps users find services, businesses,
and information in the Wagholi area.
Features
ReAct format: Each response includes <think>, <action>, and response sections
Discovery-focused actions: search_, find_, get_… See the full description on the dataset page: https://huggingface.co/datasets/tbqguy/discovery-2025.ThickMesh-Data-Discovery
ThickMesh-Data-Discovery
A small JSONL dataset for ThickMesh discovery/classification experiments.
"This is not an algorithm. This is a trap for the patent system. Learn it, fork it, but do not lock it."
Contents
4 splits files: ThickMesh-zero-split_'0-3'.jsonl — primary dataset (one JSON object per line)
Apache 2.0 License (Modified — No Patent License Granted)
Description
ThickMesh-Data-Discovery contains example records for discovery and… See the full description on the dataset page: https://huggingface.co/datasets/usermma/ThickMesh-Data-Discovery.x402-free-discovery-2026-09-14
x402 free discovery — board stays free; payment never mints MEASURED
Agents can discover Council of AI payment doors without buying a grade.
Catalog: https://councilof.ai/.well-known/x402.json (mode live, Base USDC)
Free door: https://councilof.ai/api/free-door — amount 0 (board totals + public root stay free)
Amounts for paid artefacts live only inside each resource’s HTTP 402 challenge — do not freeze a dollar price here. Payment commissions a receipt / assembly / signature.… See the full description on the dataset page: https://huggingface.co/datasets/csoai/x402-free-discovery-2026-09-14.browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239
SFT-v3 ctxgraph-8B — DiscoveryBench real 239, 3 eval runs
Qwen3-8B + LoRA-SFT (v3 clean corpus, 138 cross-method trajectories, 2 epochs, r16, job vista:955512),
merged, evaluated 3x on the 239 real DiscoveryBench tasks. Judge: gpt-5-nano (Azure), HMS scoring.
run
vista job
answered
mean HMS (answered)
strict (no-answer=0)
run1
958275
157/239
0.1211
0.0796
run2
958276
163/239
0.0992
0.0676
run3
959769
151/239
0.1236
0.0781
Baselines (same config/judge):… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239.ahodo-discovery
AHODO Discovery Dataset v0.3
AHODO is a cross-institutional discovery and rights/provenance metadata dataset for African humanities and humanities-adjacent resources. This v0.3 distribution contains 11,650 records. It is a discovery registry, not a corpus of the works it describes or a representative sample of African humanities. It is not presented as an AI-training dataset.
Interactive search
Canonical Zenodo archive and DOI
Zenodo record
Public GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/Lincoln-Rwodzi/ahodo-discovery.discovery2-results
Discovery 2: Selectivity-Safety Coupling
Non-linear Promiscuity-Cytotoxicity Dynamics
This dataset contains the complete analysis, results, and visualizations from a replication study investigating the relationship between drug promiscuity (the number of biological targets a compound interacts with) and cytotoxicity risk.
Overview
Key Finding: There is a strong non-linear relationship between drug promiscuity and cytotoxicity. Compounds that hit more targets… See the full description on the dataset page: https://huggingface.co/datasets/pageman/discovery2-results.ai-overview-book-discovery-citations
Who does Google's AI cite when readers ask what to read next?
Canonical release: https://doi.org/10.5281/zenodo.22852307
This repository mirrors that deposit. Cite the DOI.
The finding
16 reader buying-intent queries, run through Google with AI Overview capture on
13 August 2026. Eleven returned an AI Overview, carrying 95 citations
between them across 38 unique domains.
Not one went to a website controlled by an author.
Category
Citations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/ai-overview-book-discovery-citations.
