CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dataforge-labs /equity-perp-price-discovery Equity and pre-IPO perpetual prices Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets. Contents Table Record perpetual_mark_and_index_prices An instrument's mark price, index price and basis at an observation time Using the data market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.tabulartime-series-forecasting10K<n<100K1 likes2.2k downloads2h agoHugging Face02allenai /discoverybenchData-driven Discovery Benchmark from the paper: "DiscoveryBench: Towards Data-Driven Discovery with Large Language Models" 🔭 Overview DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.texttext-generationn<1K18 likes1.7k downloads1y agoHugging Face03sileod /discovery Dataset Card for Discovery Dataset Summary Discourse marker prediction with 174 markers Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure input : sentence1, sentence2, label: marker originally between sentence1 and sentence2 Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits Train/Val/Test Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/sileod/discovery.texttext-classification1M<n<10M8 likes721 downloads2y agoHugging Face04jpwahle /dblp-discovery-dataset Dataset Card for DBLP Discovery Dataset (D3) Dataset Summary DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.tabularother1M<n<10M4 likes472 downloads1y agoHugging Face05marcov /discovery_discovery_promptsourcetext1M<n<10M0 likes422 downloads2y agoHugging Face06dougalldeepmind /2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment. field value experiment LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.textn<1K0 likes354 downloads26d agoHugging Face07davidkling /hf-coding-tools-traces-discovery HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,022 query → response turns total (≈18,044 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.tabularn<1K1 likes237 downloads4mo agoHugging Face08hssling /parkinsons-evidence-to-discovery-prioritisation Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation. Dataset Summary The dataset integrates: evidence-priority scores for PD prevention and disease-modification candidates; pathway-to-intervention framework; individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.imagetabular-classificationn<1K0 likes184 downloads5mo agoHugging Face09PersonaBias /Reverse-circuit-discoverytabulartext-classification10K<n<100K0 likes181 downloads2mo agoHugging Face10hssling /pd-discovery-benchmark-dashboard Parkinson's Disease Discovery Benchmark Dashboard Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery. This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.tabularn<1K0 likes172 downloads5mo agoHugging Face11dougalldeepmind /2026-07-29-msm-philosophy-spec-focused-discovery Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript. Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c Brief finding No seed replicated. Ten seed archetypes were each run for three epochs. Under the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.textn<1K0 likes147 downloads2mo agoHugging Face12nhop /discoverybench DiscoveryBench - Alias A reformatted version of the original DiscoveryBench dataset for easier usage. 🤗 Original Dataset on HF 💻 GitHub Repository 📄 Paper (arXiv) 📁 Dataset Structure The dataset consists of real and synthetic subsets: Real Splits: real_train real_test Synthetic Splits: synth_train synth_dev synth_test Each split contains a list of tasks with references to associated CSV datasets needed to answer the query. LLMs are expected to use the… See the full description on the dataset page: https://huggingface.co/datasets/nhop/discoverybench.texttext-generation1K<n<10K0 likes136 downloads1y agoHugging Face13pkuHaowei /scaling_law_discovery_results Scaling Law Discovery Results Dataset Results dataset for the paper: "Can Language Models Discover Scaling Laws?" This dataset contains the complete collection of results from the Scaling Law Discovery (SLDBench) benchmark, where various AI agents attempt to discover mathematical scaling laws from experimental LLM training data. 🔗 Quick Links Resource Link 📄 Paper arXiv:2507.21184 📊 Original Benchmark SLDBench Dataset 🧪 Benchmark Code… See the full description on the dataset page: https://huggingface.co/datasets/pkuHaowei/scaling_law_discovery_results.textn<1K1 likes136 downloads9mo agoHugging Face14saidutta69 /red-pill-drug-discovery-formulation 🔴 RED-PILL Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language The first open instruction-tuning dataset for drug discovery & formulation development. Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions. ⚡ Quick Start from datasets import load_dataset # Load the full dataset ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.texttext-generation1K<n<10K0 likes124 downloads13d agoHugging Face15PersonaBias /Original-circuit-discoverytabulartext-classification10K<n<100K0 likes123 downloads2mo agoHugging Face16usermma /ThickMesh-Data-Discovery ThickMesh-Data-Discovery A small JSONL dataset for ThickMesh discovery/classification experiments. "This is not an algorithm. This is a trap for the patent system. Learn it, fork it, but do not lock it." Contents 4 splits files: ThickMesh-zero-split_'0-3'.jsonl — primary dataset (one JSON object per line) Apache 2.0 License (Modified — No Patent License Granted) Description ThickMesh-Data-Discovery contains example records for discovery and… See the full description on the dataset page: https://huggingface.co/datasets/usermma/ThickMesh-Data-Discovery.text1K<n<10K1 likes66 downloads4mo agoHugging Face17lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239 SFT-v3 ctxgraph-8B — DiscoveryBench real 239, 3 eval runs Qwen3-8B + LoRA-SFT (v3 clean corpus, 138 cross-method trajectories, 2 epochs, r16, job vista:955512), merged, evaluated 3x on the 239 real DiscoveryBench tasks. Judge: gpt-5-nano (Azure), HMS scoring. run vista job answered mean HMS (answered) strict (no-answer=0) run1 958275 157/239 0.1211 0.0796 run2 958276 163/239 0.0992 0.0676 run3 959769 151/239 0.1236 0.0781 Baselines (same config/judge):… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239.tabularn<1K0 likes50 downloads25d agoHugging Face18Lincoln-Rwodzi /ahodo-discovery AHODO Discovery Dataset v0.3 AHODO is a cross-institutional discovery and rights/provenance metadata dataset for African humanities and humanities-adjacent resources. This v0.3 distribution contains 11,650 records. It is a discovery registry, not a corpus of the works it describes or a representative sample of African humanities. It is not presented as an AI-training dataset. Interactive search Canonical Zenodo archive and DOI Zenodo record Public GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/Lincoln-Rwodzi/ahodo-discovery.text10K<n<100K1 likes48 downloads1d agoHugging Face19sempite /ai-overview-book-discovery-citations Who does Google's AI cite when readers ask what to read next? Canonical release: https://doi.org/10.5281/zenodo.22852307 This repository mirrors that deposit. Cite the DOI. The finding 16 reader buying-intent queries, run through Google with AI Overview capture on 13 August 2026. Eleven returned an AI Overview, carrying 95 citations between them across 38 unique domains. Not one went to a website controlled by an author. Category Citations… See the full description on the dataset page: https://huggingface.co/datasets/sempite/ai-overview-book-discovery-citations.textn<1K0 likes42 downloads6d agoHugging Face20OliverPerrin /LexiMind-Discovery LexiMind Discovery Dataset A curated multi-domain dataset for powering the LexiMind HuggingFace Space demo. Contains 1,219 items spanning academic papers, literary works, social media text, and curated technical blog posts — each annotated with topic and emotion labels. No news articles. The LexiMind model is trained on ArXiv papers and Project Gutenberg books; news data produced poor summarization results due to domain mismatch. Dataset Summary Source Type… See the full description on the dataset page: https://huggingface.co/datasets/OliverPerrin/LexiMind-Discovery.tabular1K<n<10K0 likes37 downloads7mo agoHugging Face21tasksource /discoveryhard Dataset Card for "discoveryhard" https://github.com/sileod/Discovery @inproceedings{sileo-etal-2019-mining, title = "Mining Discourse Markers for Unsupervised Sentence Representation Learning", author = "Sileo, Damien and Van De Cruys, Tim and Pradel, Camille and Muller, Philippe", booktitle = "Proceedings of the 2019 Conference of the North {A}merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/discoveryhard.text1M<n<10M2 likes34 downloads2y agoHugging Face22codezakh /gpu-forecasters-discovery-pairsCompanion artifact for GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization. Code: codezakh/gpu-surrogates. Used to evaluate whether surrogates can identify discovery moments: parent-to-child mutations where the child kernel is much faster than its parent. Each row is one parent-child kernel pair. Loading from datasets import load_dataset # all pairs ds = load_dataset("codezakh/gpu-forecasters-discovery-pairs", name="combined", split="pairs")… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/gpu-forecasters-discovery-pairs.tabular1K<n<10K0 likes26 downloads4mo agoHugging Face23tasksource /discoverybig Dataset Card for "discoverybig" More Information needed text1M<n<10M0 likes25 downloads3y agoHugging Face24dreeseaw /cleo-value-discovery Cleo Value-Discovery Benchmark A small (66-question), held-out benchmark for a failure mode that ordinary text-to-SQL evaluations miss: questions whose correct SQL depends on a literal that lives in the data, not the schema. The schema tells you a column is named status; only the data reveals its values are {'O','C','X'}. The schema shows to_date; only the data reveals that "current" is encoded as the sentinel '9999-01-01'. A one-shot text-to-SQL model has to guess these… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-value-discovery.texttable-question-answeringn<1K0 likes24 downloads4mo agoHugging Face25AI-TAX /factual-state-discovery-benchmark Factual State Discovery Benchmark Dataset for the Factual State Discovery Benchmark: Evaluating Fact Elicitation in Polish Tax Law (ACL 2026 SRW). It evaluates whether conversational agents can systematically elicit, through dialogue, all the facts of a taxpayer's situation from a real Polish tax interpretation document. Each sample pairs a factual state (a narrative of the taxpayer's situation, in Polish) with its decomposition into atomic facts — independent, verifiable claims… See the full description on the dataset page: https://huggingface.co/datasets/AI-TAX/factual-state-discovery-benchmark.tabularquestion-answeringn<1K0 likes22 downloads3mo agoHugging Face26lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2 DiscoveryBench ctxgraph-8b-clean Qwen3-8B ctxgraph fair config repeat 2/3; strict 0.0646, vista job 932507. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 157/239 answered, mean HMS 0.0983 over answered / 0.0646 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 157 Columns: 10 Columns Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2.tabularn<1K0 likes22 downloads1mo agoHugging Face27andrewinpractice /discovery-in-practice Discovery in Practice Three complete articles from https://discoveryinpractice.com/. Two are by Andrew Stewart; one Bench Tip is credited to Discovery in Practice. This is an article corpus, not a collection of experimental measurements. Preserve scientific limitations, citations, attribution, canonical links, and revision dates when reusing it. The train split is the dataset loader label; there is no evaluation split or benchmark claim. Original article content is CC BY 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/andrewinpractice/discovery-in-practice.textn<1K0 likes22 downloads19h agoHugging Face28mkita /topic-discovery-for-news-articles-testtabular100K<n<1M0 likes18 downloads9mo agoHugging Face29symbolzh /table_discoverytextn<1K0 likes18 downloads9mo agoHugging Face30fineset-io /ai-drug-discovery-papers AI for Drug Discovery Papers — FineSet A research-paper dataset on AI for Drug Discovery Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on AI for Drug Discovery Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/ai-drug-discovery-papers.tabulartext-classificationn<1K0 likes18 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.