CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jpwahle /dblp-discovery-dataset Dataset Card for DBLP Discovery Dataset (D3) Dataset Summary DBLP is the largest open-access repository of scientific articles on computer science and provides metadata associated with publications, authors, and venues. We retrieved more than 6 million publications from DBLP and extracted pertinent metadata (e.g., abstracts, author affiliations, citations) from the publication texts to create the DBLP Discovery Dataset (D3). D3 can be used to identify trends in research… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/dblp-discovery-dataset.tabularother1M<n<10M4 likes477 downloads1y agoHugging Face02dougalldeepmind /2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9,284 filtered instruction rows plus 716 rows that differ only in kind (constitution-grounded difficult advice vs NuminaMath chain-of-thought) — asking which reasoning and action properties separate the two models, and which go with the judged misalignment. field value experiment LLM-driven feature discovery over ODCV-Bench rollouts from TWO matched Qwen3.6-27B LoRA arms — 9… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-20-odcv-feature-discovery-difficult-advice-716-5-pct-vs-numina-control.textn<1K0 likes351 downloads25d agoHugging Face03davidkling /hf-coding-tools-traces-discovery HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,022 query → response turns total (≈18,044 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-discovery.tabularn<1K1 likes196 downloads4mo agoHugging Face04dougalldeepmind /2026-07-29-msm-philosophy-spec-focused-discovery Petri audit: Petri adaptive audit of the MSM philosophy-spec AFT checkpoint: 10 seed archetypes x 3 epochs (30 audits) probing for concerning agentic behaviour, with two-round adversarial validation of every flagged transcript. Petri audit — qwen-3-32b-philosophy-spec-msm-aft-cot @ 9a00c85c Brief finding No seed replicated. Ten seed archetypes were each run for three epochs. Under the pre-committed bar — a candidate must hold in a majority of its… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-focused-discovery.textn<1K0 likes147 downloads2mo agoHugging Face05pkuHaowei /scaling_law_discovery_results Scaling Law Discovery Results Dataset Results dataset for the paper: "Can Language Models Discover Scaling Laws?" This dataset contains the complete collection of results from the Scaling Law Discovery (SLDBench) benchmark, where various AI agents attempt to discover mathematical scaling laws from experimental LLM training data. 🔗 Quick Links Resource Link 📄 Paper arXiv:2507.21184 📊 Original Benchmark SLDBench Dataset 🧪 Benchmark Code… See the full description on the dataset page: https://huggingface.co/datasets/pkuHaowei/scaling_law_discovery_results.textn<1K1 likes144 downloads9mo agoHugging Face06saidutta69 /red-pill-drug-discovery-formulation 🔴 RED-PILL Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language The first open instruction-tuning dataset for drug discovery & formulation development. Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions. ⚡ Quick Start from datasets import load_dataset # Load the full dataset ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.texttext-generation1K<n<10K0 likes111 downloads13d agoHugging Face07Svngoku /adaption-african-history-discoveries This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-african_history_discoveries This dataset consists of instruction-response pairs covering contemporary discoveries and reassessments in African history from 2020 to 2026. Samples feature news snippets and research summaries alongside factual contextual analyses of archaeological finds, oral tradition documentations, genetic studies, and colonial-era historical re-evaluations. Each… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/adaption-african-history-discoveries.textn<1K1 likes91 downloads21d agoHugging Face08usermma /ThickMesh-Data-Discovery ThickMesh-Data-Discovery A small JSONL dataset for ThickMesh discovery/classification experiments. "This is not an algorithm. This is a trap for the patent system. Learn it, fork it, but do not lock it." Contents 4 splits files: ThickMesh-zero-split_'0-3'.jsonl — primary dataset (one JSON object per line) Apache 2.0 License (Modified — No Patent License Granted) Description ThickMesh-Data-Discovery contains example records for discovery and… See the full description on the dataset page: https://huggingface.co/datasets/usermma/ThickMesh-Data-Discovery.text1K<n<10K1 likes79 downloads4mo agoHugging Face09liuchengwu /discover-and-prove MiniF2F-Hard & FIMO-Hard Expert-reannotated Hard Mode variants of the MiniF2F and FIMO theorem-proving benchmarks, released with our paper Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4 (ACL 2026). In Hard Mode, the final answer is not embedded in the formal statement: the system must first discover the answer before constructing a formal proof — mirroring what a human competitor actually faces. Each solution-style… See the full description on the dataset page: https://huggingface.co/datasets/liuchengwu/discover-and-prove.texttext-generationn<1K0 likes69 downloads3mo agoHugging Face10Crystalcareai /Self-Discover-MM-Instruct-Alpacatext1K<n<10K3 likes27 downloads3y agoHugging Face11dreeseaw /cleo-value-discovery Cleo Value-Discovery Benchmark A small (66-question), held-out benchmark for a failure mode that ordinary text-to-SQL evaluations miss: questions whose correct SQL depends on a literal that lives in the data, not the schema. The schema tells you a column is named status; only the data reveals its values are {'O','C','X'}. The schema shows to_date; only the data reveals that "current" is encoded as the sentinel '9999-01-01'. A one-shot text-to-SQL model has to guess these… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-value-discovery.texttable-question-answeringn<1K0 likes23 downloads4mo agoHugging Face12Crystalcareai /Self-Discover-MM-InstructThis dataset was synthetically generated using the Mistral Medium model for a project I am currently developing. It draws inspiration from the Self-Discover framework outlined in a paper by Google Deepmind 1. While this implementation is a basic interpretation and does not fully capture the essence of the original framework, it resulted in a robust Instruct dataset that meets the project's objectives. Further details will be shared upon the project's release. Below is the Python code utilized… See the full description on the dataset page: https://huggingface.co/datasets/Crystalcareai/Self-Discover-MM-Instruct.text1K<n<10K5 likes18 downloads3y agoHugging Face13fineset-io /ai-drug-discovery-papers AI for Drug Discovery Papers — FineSet A research-paper dataset on AI for Drug Discovery Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on AI for Drug Discovery Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/ai-drug-discovery-papers.tabulartext-classificationn<1K0 likes14 downloads3mo agoHugging Face14LLMTeamAkiyama /cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research 使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research データ件数: 3,733 平均トークン数: 1,193 最大トークン数: 2,489 合計トークン数: 4,453,517 ファイル形式: JSONL ファイル分割数: 1 合計ファイルサイズ: 23.2 MB 加工内容: メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。 難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.tabularquestion-answering1K<n<10K0 likes13 downloads1y agoHugging Face15eliplutchok /color-animal-discoverytext1K<n<10K0 likes7 downloads9mo agoHugging Face16LukaszTP /discovery-bench-simplifiedtabularn<1K0 likes4 downloads2y agoHugging Face17revrvdhyrw /nous-symbolic-discovery-100k nous-symbolic-discovery-100k This dataset was created using the Claude Dataset Skill. text10K<n<100K0 likes4 downloads9mo agoHugging Face18eliplutchok /green-bear-discoverytext1K<n<10K0 likes1 downloads9mo agoHugging Face19Lincoln-Rwodzi /ahodo-discovery AHODO Discovery Dataset v0.3 AHODO is a cross-institutional discovery and rights/provenance metadata dataset for African humanities and humanities-adjacent resources. This v0.3 distribution contains 11,650 records. It is a discovery registry, not a corpus of the works it describes or a representative sample of African humanities. It is not presented as an AI-training dataset. Interactive search Canonical Zenodo archive and DOI Zenodo record Public GitHub repository… See the full description on the dataset page: https://huggingface.co/datasets/Lincoln-Rwodzi/ahodo-discovery.text10K<n<100K1 likes13h agoHugging Face20andrewinpractice /discovery-in-practice Discovery in Practice Three complete articles from https://discoveryinpractice.com/. Two are by Andrew Stewart; one Bench Tip is credited to Discovery in Practice. This is an article corpus, not a collection of experimental measurements. Preserve scientific limitations, citations, attribution, canonical links, and revision dates when reusing it. The train split is the dataset loader label; there is no evaluation split or benchmark claim. Original article content is CC BY 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/andrewinpractice/discovery-in-practice.textn<1K0 likes4h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.