CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dataforge-labs /equity-perp-price-discovery Equity and pre-IPO perpetual prices Snapshots of perpetual-futures mark prices, index prices and basis from Aevo. The instrument universe includes equities, ETFs, commodities, foreign exchange, pre-IPO contracts and crypto assets. Contents Table Record perpetual_mark_and_index_prices An instrument's mark price, index price and basis at an observation time Using the data market_type identifies the instrument category. is_rwa flags the… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/equity-perp-price-discovery.tabulartime-series-forecasting10K<n<100K1 likes2.2k downloads2h agoHugging Face02jpwahle /dblp-discovery-dataset Dataset Card for DBLP Discovery Dataset (D3) Dataset Summary DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.tabularother1M<n<10M4 likes472 downloads1y agoHugging Face03davidkling /hf-coding-tools-traces-discovery HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,022 query → response turns total (≈18,044 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.tabularn<1K1 likes237 downloads4mo agoHugging Face04hssling /parkinsons-evidence-to-discovery-prioritisation Parkinson's Disease Evidence-to-Discovery Prioritisation Dataset This Hugging Face dataset package contains processed research assets from an AI-assisted evidence synthesis and computational validation project on Parkinson's disease (PD) prevention and disease-modifying therapeutic strategy prioritisation. Dataset Summary The dataset integrates: evidence-priority scores for PD prevention and disease-modification candidates; pathway-to-intervention framework; individual… See the full description on the dataset page: https://huggingface.co/datasets/hssling/parkinsons-evidence-to-discovery-prioritisation.imagetabular-classificationn<1K0 likes184 downloads5mo agoHugging Face05PersonaBias /Reverse-circuit-discoverytabulartext-classification10K<n<100K0 likes181 downloads2mo agoHugging Face06hssling /pd-discovery-benchmark-dashboard Parkinson's Disease Discovery Benchmark Dashboard Reusable benchmark, knowledge graph, manuscript resource, and Streamlit dashboard for Parkinson's disease target-to-intervention discovery. This repository integrates evidence-synthesis priority scores, target tractability, omics/pathway recurrence, ChEMBL compound activity, RDKit physicochemical heuristics, Human Protein Atlas cell-type context, iPSC/stem-cell validation mappings, and publication-ready figures.… See the full description on the dataset page: https://huggingface.co/datasets/hssling/pd-discovery-benchmark-dashboard.tabularn<1K0 likes172 downloads5mo agoHugging Face07PersonaBias /Original-circuit-discoverytabulartext-classification10K<n<100K0 likes123 downloads2mo agoHugging Face08rocky250 /Science-Discoverytabular100K<n<1M0 likes91 downloads10mo agoHugging Face09lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239 SFT-v3 ctxgraph-8B — DiscoveryBench real 239, 3 eval runs Qwen3-8B + LoRA-SFT (v3 clean corpus, 138 cross-method trajectories, 2 epochs, r16, job vista:955512), merged, evaluated 3x on the 239 real DiscoveryBench tasks. Judge: gpt-5-nano (Azure), HMS scoring. run vista job answered mean HMS (answered) strict (no-answer=0) run1 958275 157/239 0.1211 0.0796 run2 958276 163/239 0.0992 0.0676 run3 959769 151/239 0.1236 0.0781 Baselines (same config/judge):… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-sft-v3-eval-real-239.tabularn<1K0 likes50 downloads25d agoHugging Face10OliverPerrin /LexiMind-Discovery LexiMind Discovery Dataset A curated multi-domain dataset for powering the LexiMind HuggingFace Space demo. Contains 1,219 items spanning academic papers, literary works, social media text, and curated technical blog posts — each annotated with topic and emotion labels. No news articles. The LexiMind model is trained on ArXiv papers and Project Gutenberg books; news data produced poor summarization results due to domain mismatch. Dataset Summary Source Type… See the full description on the dataset page: https://huggingface.co/datasets/OliverPerrin/LexiMind-Discovery.tabular1K<n<10K0 likes37 downloads7mo agoHugging Face11codezakh /gpu-forecasters-discovery-pairsCompanion artifact for GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization. Code: codezakh/gpu-surrogates. Used to evaluate whether surrogates can identify discovery moments: parent-to-child mutations where the child kernel is much faster than its parent. Each row is one parent-child kernel pair. Loading from datasets import load_dataset # all pairs ds = load_dataset("codezakh/gpu-forecasters-discovery-pairs", name="combined", split="pairs")… See the full description on the dataset page: https://huggingface.co/datasets/codezakh/gpu-forecasters-discovery-pairs.tabular1K<n<10K0 likes26 downloads4mo agoHugging Face12AI-TAX /factual-state-discovery-benchmark Factual State Discovery Benchmark Dataset for the Factual State Discovery Benchmark: Evaluating Fact Elicitation in Polish Tax Law (ACL 2026 SRW). It evaluates whether conversational agents can systematically elicit, through dialogue, all the facts of a taxpayer's situation from a real Polish tax interpretation document. Each sample pairs a factual state (a narrative of the taxpayer's situation, in Polish) with its decomposition into atomic facts — independent, verifiable claims… See the full description on the dataset page: https://huggingface.co/datasets/AI-TAX/factual-state-discovery-benchmark.tabularquestion-answeringn<1K0 likes22 downloads3mo agoHugging Face13lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2 DiscoveryBench ctxgraph-8b-clean Qwen3-8B ctxgraph fair config repeat 2/3; strict 0.0646, vista job 932507. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 157/239 answered, mean HMS 0.0983 over answered / 0.0646 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 157 Columns: 10 Columns Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-clean-239q-v2.tabularn<1K0 likes22 downloads1mo agoHugging Face14mkita /topic-discovery-for-news-articles-testtabular100K<n<1M0 likes18 downloads9mo agoHugging Face15fineset-io /ai-drug-discovery-papers AI for Drug Discovery Papers — FineSet A research-paper dataset on AI for Drug Discovery Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on AI for Drug Discovery Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/ai-drug-discovery-papers.tabulartext-classificationn<1K0 likes18 downloads3mo agoHugging Face16davidkling /hf-coding-tools-dashboard-discovery HuggingFace AI Coding Tools Dashboard Benchmark data from the HuggingFace AI Dashboard — tracking how AI coding tools (Claude Code, Codex, Copilot, Cursor) recommend HuggingFace products across 32 developer categories. Dataset Structure Split Description Rows results Full benchmark results with LLM responses, cost, tokens, latency, and product detection 9022 queries Benchmark query definitions across 32 categories 284 runs Run metadata and tool/model… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-dashboard-discovery.tabulartext-generation1K<n<10K0 likes17 downloads4mo agoHugging Face17LLMTeamAkiyama /cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research データ件数: 3,733 平均トークン数: 1,193 最大トークン数: 2,489 合計トークン数: 4,453,517 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 23.2 MB 加工内容: メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。 難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.tabularquestion-answering1K<n<10K0 likes16 downloads1y agoHugging Face18kenrinzero /ps1-discovery-corpus PS1 Discovery Corpus &nbsp; Canonical repo & full docs on GitHub → A multilingual metadata corpus of 7,995 PlayStation 1 games, built to be searched by feel — "a cozy fishing game with an anime aesthetic", "a 1996/97 Japanese game where you could send letters", "a bleak sci-fi adventure nobody remembers" — rather than by popularity or rigid filters. It is deliberately biased toward obscure and Japan-exclusive titles: the long tail most databases skip. The dataset's value is its… See the full description on the dataset page: https://huggingface.co/datasets/kenrinzero/ps1-discovery-corpus.tabulartext-retrieval1K<n<10K0 likes16 downloads3mo agoHugging Face19mkita /topic-discovery-for-news-articlestabular10K<n<100K0 likes14 downloads9mo agoHugging Face20lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v1 browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v1 DiscoveryBench fold-8b Qwen3-8B fold baseline repeat 1/3, original shared config; strict 0.0733, vista job 932459. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 155/239 answered, mean HMS 0.1131 over answered / 0.0733 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 155 Columns: 10 Columns Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v1.tabularn<1K0 likes13 downloads1mo agoHugging Face21lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2 DiscoveryBench ctxgraph-8b-synth synth repeat 2/3; strict 0.1250, vista job 932522. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 149/239 answered, mean HMS 0.1678 over answered / 0.1046 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 149 Columns: 10 Columns Column Type Description task_id Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-synth-239q-v2.tabularn<1K0 likes13 downloads1mo agoHugging Face22lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-v2-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-v2-239q-v2 DiscoveryBench ctxgraph-8b-dpo-v2 DPO v2 eval 2/2; strict 0.0758, answered 163, vista job 936545. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 163/239 answered, mean HMS 0.1111 over answered / 0.0758 strict-239. Part of DPO data generation: 8B strict scores 0.0846/0.0778/0.0832 (above all 30B ctxgraph runs 0.060-0.076 and on par with 30B fold 0.078-0.086); answer counts 172/177/178. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-v2-239q-v2.tabularn<1K0 likes12 downloads1mo agoHugging Face23lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-fixedops-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-fixedops-239q-v2 DiscoveryBench ctxgraph with SIX graph-op fixes (prompt example fix, junk-observation filter, explanatory op-failure feedback + eligible-id lists, auto-cleanup notices, feedback slimming) AND forced consolidation OFF (SAB_CONSOLIDATION_INTERVAL=0). vista job 928333, repeat v2. 152/239 answered, strict 0.0605, invalid-op rate 10%. Invalid-op rate down from 63% baseline; answer rate and strict score NOT… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-fixedops-239q-v2.tabularn<1K0 likes11 downloads1mo agoHugging Face24lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-239q-v3 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-239q-v3 DiscoveryBench ctxgraph-8b Qwen3-8B repeat 3 of 3, same config, vista job 932389. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 178/239 answered, mean HMS 0.1116 over answered / 0.0832 strict-239. Part of DPO data generation: 8B strict scores 0.0846/0.0778/0.0832 (above all 30B ctxgraph runs 0.060-0.076 and on par with 30B fold 0.078-0.086); answer counts 172/177/178. Dataset Info Rows: 178… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-239q-v3.tabularn<1K0 likes11 downloads1mo agoHugging Face25lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3 browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3 DiscoveryBench react-8b Qwen3-8B react baseline repeat 3/3; strict 0.0575, vista job 932458. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 132/239 answered, mean HMS 0.1041 over answered / 0.0575 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 132 Columns: 10 Columns Column Type Description task_id Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-react-8b-239q-v3.tabularn<1K0 likes11 downloads1mo agoHugging Face26lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3 browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3 DiscoveryBench fold-8b Qwen3-8B fold baseline repeat 3/3; strict 0.0728, vista job 932461. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 150/239 answered, mean HMS 0.1160 over answered / 0.0728 strict-239. Part of the 8B fair three-way + DPO data generation batch (2026-08-23/24). Dataset Info Rows: 150 Columns: 10 Columns Column Type Description task_id Value('string')… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-fold-8b-239q-v3.tabularn<1K0 likes11 downloads1mo agoHugging Face27lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2 DiscoveryBench ctxgraph-8b-dpo DPO eval 2/2; strict 0.0762, answered 156, vista job 933235. 239 queries, max_turn=24, judge gpt-5-nano (Azure). 156/239 answered, mean HMS 0.1168 over answered / 0.0762 strict-239. Part of DPO data generation: 8B strict scores 0.0846/0.0778/0.0832 (above all 30B ctxgraph runs 0.060-0.076 and on par with 30B fold 0.078-0.086); answer counts 172/177/178. Dataset Info… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-8b-dpo-239q-v2.tabularn<1K0 likes11 downloads1mo agoHugging Face28dbn4 /drug-discovery-demotabularn<1K0 likes10 downloads2mo agoHugging Face29lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-fixedops-239q-v1 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-fixedops-239q-v1 DiscoveryBench ctxgraph with SIX graph-op fixes (prompt example fix, junk-observation filter, explanatory op-failure feedback + eligible-id lists, auto-cleanup notices, feedback slimming) AND forced consolidation OFF (SAB_CONSOLIDATION_INTERVAL=0). vista job 928332, repeat v1. 164/239 answered, strict 0.0731, invalid-op rate 6%. Invalid-op rate down from 63% baseline; answer rate and strict score NOT… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-fixedops-239q-v1.tabularn<1K0 likes10 downloads1mo agoHugging Face30lingchensanwen /browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-staterecovery-239q-v1 browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-staterecovery-239q-v1 DiscoveryBench ctxgraph FULL stack: six graph-op fixes + consolidation OFF + answer-discipline prompt + STATE-RECOVERY prompt (on truncation treat code as never-run, verify state, patch only missing steps). vista job 928943, repeat v1. 139/239 answered, strict 0.0632. Invalid-op rate down from 63% baseline; answer rate and strict score NOT significantly improved vs pre-fix runs (151-157 answered… See the full description on the dataset page: https://huggingface.co/datasets/lingchensanwen/browsecomp-ctxgraph-30b-rl-discoverybench-ctxgraph-staterecovery-239q-v1.tabularn<1K0 likes10 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.