CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01trillionlabs /TheBioCollection TheBioCollection TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.texttext-generation10M<n<100M20 likes2.7k downloads2mo agoHugging Face02Trimness8 /reddit_dataset_145 Bittensor Subnet 13 Reddit Dataset Dataset Summary This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks. For more information about the dataset, please visit the official repository. Supported Tasks The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/Trimness8/reddit_dataset_145.texttext-classification10M<n<100M0 likes903 downloads2y agoHugging Face03trillionlabs /rBridge 🌉 rBridge Paper's Reasoning Traces & Token Logprobs This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks, released as part of the rBridge project (paper). rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B) can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.tabulartext-generation10K<n<100K1 likes470 downloads7mo agoHugging Face04yungisimon /data_on_trial Data on Trial — Benchmark Artifacts Benchmark artifacts for the data_on_trial jury-system pipeline. The pipeline code lives on GitHub: QuiZet/data_on_trial. These files are .gitignored in the code repo because of their size, and mirror the repository's directory layout so they can be dropped back in place. Contents Path Description datasets/ Downloaded / generated HF datasets used as pipeline inputs (Arrow/JSON). experiments/ Experiment snapshots for… See the full description on the dataset page: https://huggingface.co/datasets/yungisimon/data_on_trial.textquestion-answering1M<n<10M0 likes427 downloads2mo agoHugging Face05LingoIITGN /triveni-raw 📦 Pretraining Corpus 📊 Dataset Overview This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total — — 183,627 📁 Dataset Sources 🗣️ Vaani Dataset License: CC-BY-4.0 Description: VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/triveni-raw.imagequestion-answering100K<n<1M2 likes403 downloads1y agoHugging Face06theprint /TriviaMix TriviaMix Trivia questions generated from a mix of Wikipedia pages. Note that the answers in this dataset have not been verified. There are bound to be some errors in the data, if used as is. question-answering10K<n<100K2 likes345 downloads9mo agoHugging Face07OpenMed /Medical-Reasoning-SFT-Trinity-Mini Medical-Reasoning-SFT-Trinity-Mini A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model arcee-ai/Trinity-Mini Total Samples ~810,374 Estimated Tokens ~1.52 Billion Content Tokens ~542 Million Reasoning Tokens ~977 Million Language English Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.texttext-generation100K<n<1M79 likes226 downloads8mo agoHugging Face08LingoIITGN /Triveni 📦 Pretraining Corpus 📊 Dataset Overview This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining. Source Languages Samples per Language Total Samples Vaani Hindi, English, Hinglish 30,195 90,585 Flickr30k Hindi, English, Hinglish 31,014 93,042 Total — — 183,627 📁 Dataset Sources 🗣️ Vaani Dataset License: CC-BY-4.0 Description: VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.imagequestion-answering10K<n<100K0 likes224 downloads1y agoHugging Face09trillionlabs /TheBioCollection-Eval TheBioCollection-Eval TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets. Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.texttext-generation1K<n<10K2 likes222 downloads3mo agoHugging Face10tri-fair-lab /captrack Dataset Card for CapTrack Dataset Summary CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions: CAN (Latent Competence): What a model is capable of doing under ideal prompting WILL (Default Behavioral Preferences): What a model chooses to do by default HOW (Protocol Compliance): How reliably a… See the full description on the dataset page: https://huggingface.co/datasets/tri-fair-lab/captrack.textquestion-answering10K<n<100K1 likes188 downloads7mo agoHugging Face11avihayamor /tripmatch-ai-plan-comparisons TripMatch AI — Original vs Alternative Plan Comparisons This dataset is Amit's professor-assigned extension of TripMatch AI. It compares the original daily plan with the richer alternative plan using an LLM judge. The decision is generated by the LLM as strict JSON. Python is used only for orchestration, persistence, and JSON-schema validation; it does not calculate scores, choose a winner, or write explanations. The generation jobs use vLLM structured outputs with the published… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-plan-comparisons.tabulartext-generation10K<n<100K0 likes183 downloads28d agoHugging Face12TristanBehrens /jsfakes2024json JSFakes Chorales 2024 JSON This is a JSON representation of the dataset https://github.com/omarperacha/js-fakes. text-generation1K<n<10K1 likes170 downloads3y agoHugging Face13emgena /omnimcp_cyber_siem_triage_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cyber_siem_triage_teaser.texttext-generationn<1K0 likes159 downloads6d agoHugging Face14WillBolton /oncology-trial-strategy Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents Oncology trial-strategy decision episodes — the dataset for our ICML 2026 workshop paper. Temporal dataset for offline policy training of clinical-trial-strategy decision agents, from our ICML 2026 workshop paper, accepted at two workshops: GenBio (Generative and Agentic AI for Biology) as "Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents". Offline2Online (Decision-Making… See the full description on the dataset page: https://huggingface.co/datasets/WillBolton/oncology-trial-strategy.tabulartext-generation1K<n<10K0 likes154 downloads3mo agoHugging Face15makora-ai /triton-gpu-latency Triton GPU Latency Dataset A large dataset of PyTorch problems (mostly from KernelBench) paired with candidate Triton-kernel implementations and their measured GPU runtimes generated by MakoraGenerate. Each row is a self-contained Python program that defines (1) a reference Model written with plain PyTorch ops and (2) a ModelNew that re-implements the same forward pass with a hand-written or generated Triton kernel. The label is the runtime of executing ModelNew. Built for… See the full description on the dataset page: https://huggingface.co/datasets/makora-ai/triton-gpu-latency.texttext-generation100K<n<1M17 likes149 downloads3mo agoHugging Face16SINAI /ALIA-es-legal-administrative-triplets Dataset Introduction The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the training of embedding models and dense retrievers specialized in Spanish legal and administrative language. Hard negatives are passages that are semantically similar to a query but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.texttext-generation1M<n<10M2 likes146 downloads4mo agoHugging Face17Alberto1231 /prism_trial_3_balanced PRISM Trial 3: Fixed Balanced Cohorts This is the preregistration-ready companion to Alberto1231/prism_trial_3. Every conversation is dated 2023 or later; the observed range is November 22 through December 22, 2023. Every target is the genuine next human turn after the assistant response selected by that participant. Evaluation versus analysis Use the full configuration for model evaluation. It contains the same 456 unique held-out respondents as PRISM Trial 3, so… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/prism_trial_3_balanced.texttext-generation1K<n<10K1 likes140 downloads1mo agoHugging Face18Lots-of-LoRAs /task521_trivia_question_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task521_trivia_question_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task521_trivia_question_classification.texttext-generation1K<n<10K1 likes120 downloads2y agoHugging Face19Lots-of-LoRAs /task1565_triviaqa_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1565_triviaqa_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1565_triviaqa_classification.texttext-generationn<1K0 likes115 downloads2y agoHugging Face20DarrenLoong /TRIAGE_Bench TRIAGE-Bench: Testing Resolution of Inter-Authority Guideline Evidence TRIAGE-Bench is a benchmark for evaluating how LLMs resolve conflicts between authoritative clinical knowledge sources. It covers 2,000 items across four conflict types, each with an explicit governing policy that defines policy-consistent correctness. Key Features 2,000 benchmark items (500 per conflict type) grounded in real guideline and drug-label discrepancies Four conflict types:… See the full description on the dataset page: https://huggingface.co/datasets/DarrenLoong/TRIAGE_Bench.textquestion-answeringn<1K0 likes110 downloads2mo agoHugging Face21springofwindslabs /function-calling-ja-trial ⚡ Professional Japanese Function Calling Dataset for LLM Alignment (Free Trial) 15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%). Schema Validation Summary Programmatic validation of this exact trial file - reproducible from data.jsonl. What this trial verifies — use these 50 rows to confirm, on your own stack: Schema integrity (strict JSONL, matches the published schema) Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/function-calling-ja-trial.texttext-generationn<1K0 likes109 downloads1d agoHugging Face22springofwindslabs /regulatory-compliance-cot-trial ⚡ Regulatory Compliance & Legal CoT Dataset for Enterprise Agents (Free Trial) 15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%). Schema Validation Summary Programmatic validation of this exact trial file - reproducible from data.jsonl. What this trial verifies — use these 50 rows to confirm, on your own stack: Schema integrity (strict JSONL, matches the published schema) Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/regulatory-compliance-cot-trial.texttext-generationn<1K0 likes104 downloads1d agoHugging Face23Jack-Jieke-Wu /Paper-Writing-Exam-Trials Paper-Writing Exam Agent Trials This public dataset contains sanitized records of completed evaluations against Paper-Writing-Exam. It preserves trial summaries, event indexes, allowlisted artifacts, final submissions, and verifier results. It is not a runnable task dataset. Dataset relationship Dataset Role May it be used as a Harbor task? Paper-Writing-Exam Canonical task input Yes Paper-Writing-Exam-Trials Evidence from completed task runs No… See the full description on the dataset page: https://huggingface.co/datasets/Jack-Jieke-Wu/Paper-Writing-Exam-Trials.text-generation0 likes100 downloads20d agoHugging Face24springofwindslabs /function-calling-en-trial ⚡ Professional Function Calling Dataset for LLM Alignment (Free Trial) 15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%). Schema Validation Summary Programmatic validation of this exact trial file - reproducible from data.jsonl. What this trial verifies — use these 50 rows to confirm, on your own stack: Schema integrity (strict JSONL, matches the published schema) Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/function-calling-en-trial.texttext-generationn<1K0 likes100 downloads1d agoHugging Face25amd /InstructGpt-TriviaQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.texttext-generation1M<n<10M0 likes98 downloads7mo agoHugging Face26Bryan35406 /agi-trilogy AGI Trilogy · AGI 삼부작 License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md. Three short stories and a companion piece on the arrival of AGI — three countries, three tenses that do not translate into one another. Co-written by a human author and a large language model, in Korean and English as a scene-aligned mirror. Published at… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/agi-trilogy.texttext-generationn<1K0 likes98 downloads1mo agoHugging Face27navimusaget /theogonos-trilogy Theogonos Trilogy — AI-Readable Literary Corpus Author: Navi MusagetLanguage: EnglishLicense: Creative Commons Attribution 4.0 International (CC BY 4.0)Pattern: 7-3-1-8 This repository hosts the complete machine-readable text corpus of the Theogonos Trilogy, a philosophical science-fiction sequence by Navi Musaget. The trilogy consists of: Lunar Bell Protocol Theogonos: The Primordial Code Protocol Amor The corpus is intended for literary analysis, philosophy of intelligence, AI… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-trilogy.texttext-generationn<1K0 likes97 downloads4mo agoHugging Face28Sigurdur /is-trivia-questions Icelandic trivia questions Icelandic trivia question compiled and created by Sveinn Steinarsson, Valur Freyr Steinarsson, and Svavar Kjarrval https://github.com/sveinn-steinarsson/is-trivia-questions Dálkanúmer Valfrjálst Lýsing 1 Nei Flokkanúmer 2 Já Undirflokkur ef til staðar 3 Nei Erfiðleikastig: 1: Létt, 2: Meðal, 3: Erfið 4 Já Gæðastig: 1: Slöpp, 2: Góð, 3: Ágæt 5 Nei Spurningin 6 Nei Svarið Flokkanúmer Flokkanafn 1 Almenn kunnátta 2 Náttúra… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/is-trivia-questions.tabularquestion-answering10K<n<100K0 likes95 downloads2y agoHugging Face29fhai50032 /latentsig-med-triage-router LatentSig Medical Triage Router Dataset 1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers. Overview This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must: Select the correct tool from 7 available medical tools Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.imagetext-generation1K<n<10K0 likes93 downloads4mo agoHugging Face30legalnlpresearcher /legal-statutory-triage-sft Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860). 1. Overview and Scope With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.texttext-generation1K<n<10K0 likes91 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.