CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01trillionlabs /TheBioCollection TheBioCollection TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.texttext-generation10M<n<100M20 likes2.7k downloads2mo agoHugging Face02Trimness8 /reddit_dataset_145 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/Trimness8/reddit_dataset_145.texttext-classification10M<n<100M0 likes876 downloads2y agoHugging Face03yungisimon /data_on_trial Data on Trial — Benchmark Artifacts Benchmark artifacts for the data_on_trial jury-system pipeline. The pipeline code lives on GitHub: QuiZet/data_on_trial. These files are .gitignored in the code repo because of their size, and mirror the repository's directory layout so they can be dropped back in place. Contents Path Description datasets/ Downloaded / generated HF datasets used as pipeline inputs (Arrow/JSON). experiments/ Experiment snapshots for… See the full description on the dataset page: https://huggingface.co/datasets/yungisimon/data_on_trial.textquestion-answering1M<n<10M0 likes426 downloads3mo agoHugging Face04LingoIITGN /triveni-raw 📦 Pretraining Corpus 📊 Dataset Overview This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total — — 183,627 📁 Dataset Sources 🗣️ Vaani Dataset License: CC-BY-4.0 Description: VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/triveni-raw.imagequestion-answering100K<n<1M2 likes402 downloads1y agoHugging Face05trillionlabs /rBridge 🌉 rBridge Paper's Reasoning Traces & Token Logprobs This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks, released as part of the rBridge project (paper). rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B) can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.tabulartext-generation10K<n<100K1 likes362 downloads7mo agoHugging Face06OpenMed /Medical-Reasoning-SFT-Trinity-Mini Medical-Reasoning-SFT-Trinity-Mini A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model arcee-ai/Trinity-Mini Total Samples ~810,374 Estimated Tokens ~1.52 Billion Content Tokens ~542 Million Reasoning Tokens ~977 Million Language English Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.texttext-generation100K<n<1M79 likes250 downloads8mo agoHugging Face07trillionlabs /TheBioCollection-Eval TheBioCollection-Eval TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets. Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.texttext-generation1K<n<10K2 likes232 downloads3mo agoHugging Face08LingoIITGN /Triveni 📦 Pretraining Corpus 📊 Dataset Overview This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total — — 183,627 📁 Dataset Sources 🗣️ Vaani Dataset License: CC-BY-4.0 Description: VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.imagequestion-answering10K<n<100K0 likes227 downloads1y agoHugging Face09tri-fair-lab /captrack Dataset Card for CapTrack Dataset Summary CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions: CAN (Latent Competence): What a model is capable of doing under ideal prompting WILL (Default Behavioral Preferences): What a model chooses to do by default HOW (Protocol Compliance): How reliably a… See the full description on the dataset page: https://huggingface.co/datasets/tri-fair-lab/captrack.textquestion-answering10K<n<100K1 likes175 downloads7mo agoHugging Face10avihayamor /tripmatch-ai-plan-comparisons TripMatch AI — Original vs Alternative Plan Comparisons This dataset is Amit's professor-assigned extension of TripMatch AI. It compares the original daily plan with the richer alternative plan using an LLM judge. The decision is generated by the LLM as strict JSON. Python is used only for orchestration, persistence, and JSON-schema validation; it does not calculate scores, choose a winner, or write explanations. The generation jobs use vLLM structured outputs with the published… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-plan-comparisons.tabulartext-generation10K<n<100K0 likes171 downloads29d agoHugging Face11WillBolton /oncology-trial-strategy Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents Oncology trial-strategy decision episodes — the dataset for our ICML 2026 workshop paper. Temporal dataset for offline policy training of clinical-trial-strategy decision agents, from our ICML 2026 workshop paper, accepted at two workshops: GenBio (Generative and Agentic AI for Biology) as "Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents". Offline2Online (Decision-Making… See the full description on the dataset page: https://huggingface.co/datasets/WillBolton/oncology-trial-strategy.tabulartext-generation1K<n<10K0 likes160 downloads3mo agoHugging Face12emgena /omnimcp_cyber_siem_triage_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cyber_siem_triage_teaser.texttext-generationn<1K0 likes160 downloads8d agoHugging Face13makora-ai /triton-gpu-latency Triton GPU Latency Dataset A large dataset of PyTorch problems (mostly from KernelBench) paired with candidate Triton-kernel implementations and their measured GPU runtimes generated by MakoraGenerate. Each row is a self-contained Python program that defines (1) a reference Model written with plain PyTorch ops and (2) a ModelNew that re-implements the same forward pass with a hand-written or generated Triton kernel. The label is the runtime of executing ModelNew. Built for… See the full description on the dataset page: https://huggingface.co/datasets/makora-ai/triton-gpu-latency.texttext-generation100K<n<1M17 likes159 downloads3mo agoHugging Face14SINAI /ALIA-es-legal-administrative-triplets Dataset Introduction The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish legal and administrative language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.texttext-generation1M<n<10M2 likes142 downloads4mo agoHugging Face15Alberto1231 /prism_trial_3_balanced PRISM Trial 3: Fixed Balanced Cohorts This is the preregistration-ready companion to Alberto1231/prism_trial_3. Every conversation is dated 2023 or later; the observed range is November 22 through December 22, 2023. Every target is the genuine next human turn after the assistant response selected by that participant. Evaluation versus analysis Use the full configuration for model evaluation. It contains the same 456 unique held-out respondents as PRISM Trial 3, so… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/prism_trial_3_balanced.texttext-generation1K<n<10K1 likes138 downloads2mo agoHugging Face16springofwindslabs /function-calling-ja-trial ⚡ Professional Japanese Function Calling Dataset for LLM Alignment (Free Trial) Official Open‑Source Evaluation Package (50 Rows Subset) by springofwindslabs 👉 Looking for full production data? The complete 1,000-row standard volume and 2,300+ row mutually exclusive, non-overlapping extended package (Total 3,300+ unique rows) are fully available for commercial deployment via our official procurement gateway: ➔… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/function-calling-ja-trial.texttext-generationn<1K0 likes133 downloads1d agoHugging Face17springofwindslabs /regulatory-compliance-cot-trial ⚡ Regulatory Compliance & Legal CoT Dataset for Enterprise Agents (Free Trial) Official Open‑Source Evaluation Package (50 Rows Subset) by springofwindslabs 👉 Looking for full production data? The complete 1,000-row standard volume and 2,300+ row mutually exclusive, non-overlapping extended package (Total 3,300+ unique rows) are fully available for commercial deployment via our official procurement gateway: ➔… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/regulatory-compliance-cot-trial.texttext-generationn<1K0 likes126 downloads1d agoHugging Face18springofwindslabs /function-calling-en-trial ⚡ Professional Function Calling Dataset for LLM Alignment (Free Trial) Official Open‑Source Evaluation Package (50 Rows Subset) by springofwindslabs 👉 Looking for full production data? The complete 1,000-row standard volume and 2,300+ row mutually exclusive, non-overlapping extended package (Total 3,300+ unique rows) are fully available for commercial deployment via our official procurement gateway: ➔… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/function-calling-en-trial.texttext-generationn<1K0 likes120 downloads1d agoHugging Face19TimotheeB /triage-medical-dataset Dataset release Version: 2026-03-19-v1 Published at: 2026-03-19T15:42:16+00:00 Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset Dataset Card - POC Triage Medical Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance. 1) Description Dataset bilingue FR/EN pour triage medical initial. Le pipeline produit deux artefacts principaux: SFT: paires instruction/reponse pour le fine-tuning supervise. DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.tabulartext-generation10K<n<100K0 likes117 downloads6mo agoHugging Face20DarrenLoong /TRIAGE_Bench TRIAGE-Bench: Testing Resolution of Inter-Authority Guideline Evidence TRIAGE-Bench is a benchmark for evaluating how LLMs resolve conflicts between authoritative clinical knowledge sources. It covers 2,000 items across four conflict types, each with an explicit governing policy that defines policy-consistent correctness. Key Features 2,000 benchmark items (500 per conflict type) grounded in real guideline and drug-label discrepancies Four conflict types:… See the full description on the dataset page: https://huggingface.co/datasets/DarrenLoong/TRIAGE_Bench.textquestion-answeringn<1K0 likes116 downloads2mo agoHugging Face21Lots-of-LoRAs /task521_trivia_question_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task521_trivia_question_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task521_trivia_question_classification.texttext-generation1K<n<10K1 likes112 downloads2y agoHugging Face22Bryan35406 /agi-trilogy AGI Trilogy · AGI 삼부작 License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md. Three short stories and a companion piece on the arrival of AGI — three countries, three tenses that do not translate into one another. Co-written by a human author and a large language model, in Korean and English as a scene-aligned mirror. Published at… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/agi-trilogy.texttext-generationn<1K0 likes109 downloads1mo agoHugging Face23fhai50032 /latentsig-med-triage-router LatentSig Medical Triage Router Dataset 1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers. Overview This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must: Select the correct tool from 7 available medical tools Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.imagetext-generation1K<n<10K0 likes103 downloads4mo agoHugging Face24Lots-of-LoRAs /task1565_triviaqa_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1565_triviaqa_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1565_triviaqa_classification.texttext-generationn<1K0 likes96 downloads2y agoHugging Face25amd /InstructGpt-TriviaQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.texttext-generation1M<n<10M0 likes96 downloads7mo agoHugging Face26Sigurdur /is-trivia-questions Icelandic trivia questions Icelandic trivia question compiled and created by Sveinn Steinarsson, Valur Freyr Steinarsson, and Svavar Kjarrval https://github.com/sveinn-steinarsson/is-trivia-questions Dálkanúmer Valfrjálst Lýsing 1 Nei Flokkanúmer 2 Já Undirflokkur ef til staðar 3 Nei Erfiðleikastig: 1: Létt, 2: Meðal, 3: Erfið 4 Já Gæðastig: 1: Slöpp, 2: Góð, 3: Ágæt 5 Nei Spurningin 6 Nei Svarið Flokkanúmer Flokkanafn 1 Almenn kunnátta 2 Náttúra… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/is-trivia-questions.tabularquestion-answering10K<n<100K0 likes94 downloads2y agoHugging Face27legalnlpresearcher /legal-statutory-triage-sft Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860). 1. Overview and Scope With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.texttext-generation1K<n<10K0 likes94 downloads25d agoHugging Face28navimusaget /theogonos-trilogy Theogonos Trilogy — AI-Readable Literary Corpus Author: Navi MusagetLanguage: EnglishLicense: Creative Commons Attribution 4.0 International (CC BY 4.0)Pattern: 7-3-1-8 This repository hosts the complete machine-readable text corpus of the Theogonos Trilogy, a philosophical science-fiction sequence by Navi Musaget. The trilogy consists of: Lunar Bell Protocol Theogonos: The Primordial Code Protocol Amor The corpus is intended for literary analysis, philosophy of intelligence, AI… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-trilogy.texttext-generationn<1K0 likes92 downloads4mo agoHugging Face29trillionlabs /SimScholar-SFT S3 SFT Trajectories Complete ReAct trajectories for scientific-literature search. Code · S3 collection · Source corpus This dataset contains 14,633 single- and two-hop tool-use trajectories. In each trajectory, a policy searches and reads a fixed scientific corpus through nine tools, then submits an answer with a correctness label. The messages column uses OpenAI tool-calling chat format. At a glance Question type Rows Correct Incorrect Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.tabularquestion-answering10K<n<100K0 likes84 downloads2mo agoHugging Face30beatsprom /cuda-triton-gpu-kernels-2026 ⚡ Complete 2026 CUDA & OpenAI Triton High-Performance GPU Kernel Engineering SFT/DPO Suite The definitive, production-grade synthetic alignment dataset engineered for training and fine-tuning open-weights Large Language Models (Qwen 2.5 Coder, DeepSeek-Coder, Llama 3.1) on ultra-high-throughput GPU kernel programming: NVIDIA Hopper H100 / Blackwell B200 TMA async transfers, OpenAI Triton 3.1+ FlashAttention-3, 32-bank conflict elimination, and low-bit FP8 / INT4 GEMM… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cuda-triton-gpu-kernels-2026.tabulartext-generation1K<n<10K0 likes83 downloads14d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.