CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01poolside-laguna-hackathon /protein-ligand-design 🧪 Protein-Ligand Design Gym — Team JAMMY poolside Laguna Hackathon submission. A tool-use reinforcement-learning environment that teaches an LLM to reason like a bench computational chemist / protein engineer — by measuring, not guessing. The problem Proteins are the molecular machines inside living cells, each built from a long string of amino-acid "letters". Ligands are the small molecules — most drugs among them — that bind to a protein to switch it on or… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/protein-ligand-design.textquestion-answering1K<n<10K1 likes205 downloads3mo agoHugging Face02build-small-hackathon /figment-eval-traces Figment Eval Traces Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders. These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment. Dataset Summary The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.tabulartext-generation100K<n<1M0 likes48 downloads3mo agoHugging Face03somosnlp-hackathon-2025 /ibero-tales-es Conjunto de datos de historias sintéticas de mitos y leyendas iberoamericanos. ⚠️ Este dataset se encuentra en desarrollo activo. Se planea expandir significativamente el número de registros y refinar el proceso de generación para mejorar la diversidad y calidad narrativa. 📚 Descripción Dataset de historias sintéticas generadas a partir de mitos y leyendas de Iberoamérica, curado y estructurado para el entrenamiento y alineamiento de modelos de lenguaje en narrativa… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/ibero-tales-es.textquestion-answering1K<n<10K0 likes26 downloads1y agoHugging Face04somosnlp-hackathon-2025 /ec-prompts-refranestextquestion-answeringn<1K0 likes25 downloads1y agoHugging Face05somosnlp-hackathon-2025 /cenia-team-sabiduriapopulartextquestion-answeringn<1K0 likes24 downloads1y agoHugging Face06lablab-ai-amd-developer-hackathon /OncoAgent-Clinical-266K 🧬 OncoAgent Clinical Dataset — 266K Curated Multi-Source Oncology Training Dataset AMD Developer Hackathon 2026 · Used to fine-tune OncoAgent v1.0 Dataset Description This dataset contains 266,854 clinical oncology training samples curated for fine-tuning large language models on cancer diagnosis, treatment recommendation, and clinical reasoning tasks. Composition Source Samples Description PMC-Patients ~100,000 Real clinical case presentations… See the full description on the dataset page: https://huggingface.co/datasets/lablab-ai-amd-developer-hackathon/OncoAgent-Clinical-266K.text-generation100K<n<1M6 likes21 downloads5mo agoHugging Face07somosnlp-hackathon-2026 /Onexe-QA-Dataset Dataset Card: Canarian Linguistic Evaluation Dataset (QA without Answers) Dataset Summary This dataset has been designed specifically for evaluating the dialectal, linguistic, and cultural understanding of Large Language Models (LLMs) within the context of Canarian Spanish. It contains 4,683 evaluation questions based on the official lexicon of the Academy of Canarian Language (Academia Canaria de la Lengua - ACL). Each record presents a linguistic query phrased… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2026/Onexe-QA-Dataset.textquestion-answering1K<n<10K0 likes19 downloads3mo agoHugging Face08somosnlp-hackathon-2025 /exam_zh_multitopic_dialect_culture exam_zh_multitopic_dialect_culture This dataset contains 300 multiple-choice questions (MCQs) from a variety of Mandarin-based assessments, spanning both regional dialect comprehension and cultural/general knowledge. 📚 Description The questions come from publicly available Chinese-language exams and quizzes, and fall into two major categories: 🗣️ Regional Dialect Tests These assess language understanding across major Chinese dialects and topolects: Hakka… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/exam_zh_multitopic_dialect_culture.tabularmultiple-choicen<1K0 likes15 downloads1y agoHugging Face09build-small-hackathon /PaperProf-traces PaperProf Agent Trace Step-by-step trace of PaperProf, an AI study buddy that turns course PDFs into interactive quiz sessions. What's in this dataset Each row in paperprof_trace.jsonl is one LLM call. Fields: Field Description session_id Groups steps from the same session step Step index within the session (1–4) type question_generation / answer_evaluation / mcq_generation topic Domain of the source chunk input Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.tabularquestion-answeringn<1K0 likes15 downloads3mo agoHugging Face10build-small-hackathon /fabella-traces Fabella Anonymized Agent Traces A public, anonymized log of the LangGraph ReAct loop inside Fabella, a small-model Gradio Space for parents who need help explaining hard things to their child in kid-appropriate language. The dataset exists for the Sharing is Caring merit badge in the Build Small Hackathon. The first version of every explanation is drafted by google/gemma-4-E4B-it via a LangGraph ReAct loop with one tool (validate_explanation). A second small model —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/fabella-traces.tabulartext-generationn<1K0 likes11 downloads3mo agoHugging Face11build-small-hackathon /packetcourt-golden-cases PacketCourt Golden Cases A small evidence-first evaluation set for auditing front-of-pack claims against the text printed on the same Indian packaged-food label. Each record contains: front-label claim text back-label evidence text expected claims and conservative verdicts expected persuasion-gap concepts expected deterministic date or whole-packet calculations The initial set is intentionally small and hand-audited. It is a regression and demonstration asset, not a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/packetcourt-golden-cases.texttext-classificationn<1K0 likes10 downloads3mo agoHugging Face12build-small-hackathon /agent-eval-golden-dataset Tech Interview Agent — Golden Eval Dataset Stop guessing whether your AI interviewer is good. Start measuring it. This dataset provides ground-truth benchmarks for evaluating AI agents that conduct tech job interviews. Each record is a structured test case: give it to your agent, collect the response, run it through the AI Agent Evaluation Pipeline, and get objective scores — no human review needed. Generated by NVIDIA Nemotron-3-Nano-30B-A3B. What's inside 40… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agent-eval-golden-dataset.text-generationn<1K0 likes8 downloads4mo agoHugging Face13pransfries /llm-hackathontext-classification1K<n<10K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.