CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saidutta69 /red-pill-drug-discovery-formulation 🔴 RED-PILL Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language The first open instruction-tuning dataset for drug discovery & formulation development. Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions. ⚡ Quick Start from datasets import load_dataset # Load the full dataset ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.texttext-generation1K<n<10K0 likes111 downloads12d agoHugging Face02liuchengwu /discover-and-prove MiniF2F-Hard & FIMO-Hard Expert-reannotated Hard Mode variants of the MiniF2F and FIMO theorem-proving benchmarks, released with our paper Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4 (ACL 2026). In Hard Mode, the final answer is not embedded in the formal statement: the system must first discover the answer before constructing a formal proof — mirroring what a human competitor actually faces. Each solution-style… See the full description on the dataset page: https://huggingface.co/datasets/liuchengwu/discover-and-prove.texttext-generationn<1K0 likes69 downloads3mo agoHugging Face03Dans-DiscountModels /RUCAIBox-Story-Generation-Alpacahttps://huggingface.co/datasets/RUCAIBox/Story-Generation RUC AI Box HC Story Generation augmented and converted to alpaca format. No filtering has been done. texttext-generation1K<n<10K13 likes64 downloads3y agoHugging Face04Banodoco /discord-archive Discord Archive This is an archive of messages from the Banodoco Discord community, where technical and artistic practitioners have been discussing open source AI art for the past three years. The archive captures a long-running community record of people learning, training, evaluating, and using open source AI art models in practice. It contains discussion around model releases, workflows, tooling, troubleshooting, creative experiments, training details, and the many small… See the full description on the dataset page: https://huggingface.co/datasets/Banodoco/discord-archive.tabulartext-generation1M<n<10M4 likes64 downloads4mo agoHugging Face05DiscoPosse /cc-traces-weka-with-subagents-051826 CC Traces — Weka, With Subagents, v5 only (May 18 2026) A collection of 96 multi-turn agentic traces drawn from real production traffic against the Claude Code CLI ≥ 2.1.139. Each trace captures the full request/response sequence of a single agent session, including per-request KV block hashes AND the original sub-agent fan-out structure (Task-tool spawned sub-agents grouped into WekaSubagentEntry blocks). With-subagents, v5-only variant. Companion to… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/cc-traces-weka-with-subagents-051826.texttext-generationn<1K0 likes57 downloads4mo agoHugging Face06Dans-DiscountModels /Alpaca_Evol_Instruct_CleanedAlpaca Evol Instruct cleaned of refusals, scrubbed of overly repetitive responses, aggresively deduplicated, and all URLs removed from the output. The final dataset has aproximately 54k instructions. Base dataset https://huggingface.co/datasets/victor123/evol_instruct_70k texttext-generation100K<n<1M6 likes52 downloads3y agoHugging Face07SINAI /ALIA-es-discriminative-stance-detection Dataset Introduction This corpus comprises 3,000 manually annotated instances for stance detection in Spanish, built from real citizen comments posted on the Decide Madrid participatory democracy platform. Each instance consists of a civic topic (target) — defined by its title and description — paired with a citizen comment, annotated for stance as favor, against, or neutral by 3 independent human annotators. The dataset is published in full accordance with the principles of… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-stance-detection.texttext-classification1K<n<10K0 likes50 downloads4mo agoHugging Face08AmareshHebbar /discharge-qa-sft Discharge Summary Q&A Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Discharge summaries + questions → precise clinical answers Why download this Build systems that answer specific questions about a patient's hospitalization from their discharge summary. Key for patient safety, care transitions, and clinical… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/discharge-qa-sft.texttext-generation10K<n<100K0 likes41 downloads3mo agoHugging Face09DiscoPosse /RAGPulse RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems 🌐 Github Link | 🤗 Workload Trace | 📑 Arxiv Paper | 🤖 How to use? RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DiscoPosse/RAGPulse.tabulartext-generation1K<n<10K0 likes30 downloads4mo agoHugging Face10CoShin /discrete_prompting_webqsp WebQSP Verbalized This dataset is derived from the WebQSP benchmark and extended with multiple graph-to-text verbalization strategies.It is designed to evaluate how different natural language representations of knowledge graphs affect large language models in knowledge-augmented QA tasks. Dataset Structure Splits: train, validation, test Format: JSONL (one JSON object per line) textquestion-answering1K<n<10K0 likes28 downloads1y agoHugging Face11dreeseaw /cleo-value-discovery Cleo Value-Discovery Benchmark A small (66-question), held-out benchmark for a failure mode that ordinary text-to-SQL evaluations miss: questions whose correct SQL depends on a literal that lives in the data, not the schema. The schema tells you a column is named status; only the data reveals its values are {'O','C','X'}. The schema shows to_date; only the data reveals that "current" is encoded as the sentinel '9999-01-01'. A one-shot text-to-SQL model has to guess these… See the full description on the dataset page: https://huggingface.co/datasets/dreeseaw/cleo-value-discovery.texttable-question-answeringn<1K0 likes23 downloads4mo agoHugging Face12theprint /databird-discoursetexttext-generation1K<n<10K0 likes17 downloads1y agoHugging Face13Dans-DiscountModels /text-splitter-alpacahttps://huggingface.co/datasets/mhenrichsen/context-aware-splits-english texttext-generation10K<n<100K1 likes15 downloads3y agoHugging Face14fineset-io /ai-drug-discovery-papers AI for Drug Discovery Papers — FineSet A research-paper dataset on AI for Drug Discovery Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on AI for Drug Discovery Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/ai-drug-discovery-papers.tabulartext-classificationn<1K0 likes14 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.