datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TheBioCollection
TheBioCollection
TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.reddit_dataset_145
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/Trimness8/reddit_dataset_145.rBridge
🌉 rBridge Paper's Reasoning Traces & Token Logprobs
This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks,
released as part of the rBridge project
(paper).
rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood
over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B)
can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.data_on_trial
Data on Trial — Benchmark Artifacts
Benchmark artifacts for the data_on_trial jury-system pipeline. The pipeline code
lives on GitHub: QuiZet/data_on_trial.
These files are .gitignored in the code repo because of their size, and mirror the
repository's directory layout so they can be dropped back in place.
Contents
Path
Description
datasets/
Downloaded / generated HF datasets used as pipeline inputs (Arrow/JSON).
experiments/
Experiment snapshots for… See the full description on the dataset page: https://huggingface.co/datasets/yungisimon/data_on_trial.triveni-raw
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/triveni-raw.TriviaMix
TriviaMix
Trivia questions generated from a mix of Wikipedia pages.
Note that the answers in this dataset have not been verified. There are bound to be some errors in the data, if used as is.
Medical-Reasoning-SFT-Trinity-Mini
Medical-Reasoning-SFT-Trinity-Mini
A large-scale medical reasoning dataset generated using arcee-ai/Trinity-Mini, containing over 810,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
arcee-ai/Trinity-Mini
Total Samples
~810,374
Estimated Tokens
~1.52 Billion
Content Tokens
~542 Million
Reasoning Tokens
~977 Million
Language
English
Schema
Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Trinity-Mini.Triveni
📦 Pretraining Corpus
📊 Dataset Overview
This dataset combines data from two major sources—Vaani and Flickr30k—to support multilingual and multimodal model pretraining.
Source
Languages
Samples per Language
Total Samples
Vaani
Hindi, English, Hinglish
30,195
90,585
Flickr30k
Hindi, English, Hinglish
31,014
93,042
Total
—
—
183,627
📁 Dataset Sources
🗣️ Vaani Dataset
License: CC-BY-4.0
Description:
VAANI is an… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/Triveni.TheBioCollection-Eval
TheBioCollection-Eval
TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets.
Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.captrack
Dataset Card for CapTrack
Dataset Summary
CapTrack is a comprehensive evaluation suite designed to measure capability drift and forgetting in Large Language Models (LLMs). The dataset enables systematic assessment of model behavior across three complementary dimensions:
CAN (Latent Competence): What a model is capable of doing under ideal prompting
WILL (Default Behavioral Preferences): What a model chooses to do by default
HOW (Protocol Compliance): How reliably a… See the full description on the dataset page: https://huggingface.co/datasets/tri-fair-lab/captrack.tripmatch-ai-plan-comparisons
TripMatch AI — Original vs Alternative Plan Comparisons
This dataset is Amit's professor-assigned extension of TripMatch AI. It compares
the original daily plan with the richer alternative plan using an LLM judge.
The decision is generated by the LLM as strict JSON. Python is used only for
orchestration, persistence, and JSON-schema validation; it does not calculate
scores, choose a winner, or write explanations.
The generation jobs use vLLM structured outputs with the published… See the full description on the dataset page: https://huggingface.co/datasets/avihayamor/tripmatch-ai-plan-comparisons.jsfakes2024json
JSFakes Chorales 2024 JSON
This is a JSON representation of the dataset https://github.com/omarperacha/js-fakes.
omnimcp_cyber_siem_triage_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cyber_siem_triage_teaser.oncology-trial-strategy
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Oncology trial-strategy decision episodes — the dataset for our ICML 2026 workshop paper.
Temporal dataset for offline policy training of clinical-trial-strategy decision agents, from our
ICML 2026 workshop paper, accepted at two workshops:
GenBio (Generative and Agentic AI for Biology) as "Learning Clinical-Trial Strategy: Offline
Policy Training for Decision Agents".
Offline2Online (Decision-Making… See the full description on the dataset page: https://huggingface.co/datasets/WillBolton/oncology-trial-strategy.triton-gpu-latency
Triton GPU Latency Dataset
A large dataset of PyTorch problems (mostly from KernelBench) paired with candidate Triton-kernel implementations and their measured GPU runtimes generated by MakoraGenerate. Each row is a self-contained Python program that defines (1) a reference Model written with plain PyTorch ops and (2) a ModelNew that re-implements the same forward pass with a hand-written or generated Triton kernel. The label is the runtime of executing ModelNew.
Built for… See the full description on the dataset page: https://huggingface.co/datasets/makora-ai/triton-gpu-latency.ALIA-es-legal-administrative-triplets
Dataset Introduction
The dataset ALIA Spanish Legal and Administrative Triplets Corpus contains hard negatives for dense retrieval training
generated from <query, passage> pairs contained in SINAI/ALIA-es-legal-administrative-triplets.The dataset was created as part of the ALIA project to improve the
training of embedding models and dense retrievers specialized in Spanish
legal and administrative language.
Hard negatives are passages that are semantically similar to a query
but not… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-legal-administrative-triplets.prism_trial_3_balanced
PRISM Trial 3: Fixed Balanced Cohorts
This is the preregistration-ready companion to Alberto1231/prism_trial_3.
Every conversation is dated 2023 or later; the observed range is November
22 through December 22, 2023. Every target is the genuine next human turn after
the assistant response selected by that participant.
Evaluation versus analysis
Use the full configuration for model evaluation. It contains the same 456
unique held-out respondents as PRISM Trial 3, so… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/prism_trial_3_balanced.task521_trivia_question_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task521_trivia_question_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task521_trivia_question_classification.task1565_triviaqa_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1565_triviaqa_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1565_triviaqa_classification.TRIAGE_Bench
TRIAGE-Bench: Testing Resolution of Inter-Authority Guideline Evidence
TRIAGE-Bench is a benchmark for evaluating how LLMs resolve conflicts between authoritative clinical knowledge sources. It covers 2,000 items across four conflict types, each with an explicit governing policy that defines policy-consistent correctness.
Key Features
2,000 benchmark items (500 per conflict type) grounded in real guideline and drug-label discrepancies
Four conflict types:… See the full description on the dataset page: https://huggingface.co/datasets/DarrenLoong/TRIAGE_Bench.function-calling-ja-trial
⚡ Professional Japanese Function Calling Dataset for LLM Alignment (Free Trial)
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/function-calling-ja-trial.regulatory-compliance-cot-trial
⚡ Regulatory Compliance & Legal CoT Dataset for Enterprise Agents (Free Trial)
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/regulatory-compliance-cot-trial.Paper-Writing-Exam-Trials
Paper-Writing Exam Agent Trials
This public dataset contains sanitized records of completed evaluations against
Paper-Writing-Exam.
It preserves trial summaries, event indexes, allowlisted artifacts, final
submissions, and verifier results. It is not a runnable task dataset.
Dataset relationship
Dataset
Role
May it be used as a Harbor task?
Paper-Writing-Exam
Canonical task input
Yes
Paper-Writing-Exam-Trials
Evidence from completed task runs
No… See the full description on the dataset page: https://huggingface.co/datasets/Jack-Jieke-Wu/Paper-Writing-Exam-Trials.function-calling-en-trial
⚡ Professional Function Calling Dataset for LLM Alignment (Free Trial)
15-second demo: strict JSONL trajectories + 7-point rubric validation (schema stability 100%).
Schema Validation Summary
Programmatic validation of this exact trial file - reproducible from data.jsonl.
What this trial verifies — use these 50 rows to confirm, on your own stack:
Schema integrity (strict JSONL, matches the published schema)
Multi-turn / tool-use structural consistency… See the full description on the dataset page: https://huggingface.co/datasets/springofwindslabs/function-calling-en-trial.InstructGpt-TriviaQa
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-TriviaQa.agi-trilogy
AGI Trilogy · AGI 삼부작
License note for ML practitioners: use of this dataset for machine learning and AI model training is expressly permitted — no further permission needed. All other rights reserved. Full terms: NOTICE.md.
Three short stories and a companion piece on the arrival of AGI — three countries, three tenses that do not translate into one another. Co-written by a human author and a large language model, in Korean and English as a scene-aligned mirror. Published at… See the full description on the dataset page: https://huggingface.co/datasets/Bryan35406/agi-trilogy.theogonos-trilogy
Theogonos Trilogy — AI-Readable Literary Corpus
Author: Navi MusagetLanguage: EnglishLicense: Creative Commons Attribution 4.0 International (CC BY 4.0)Pattern: 7-3-1-8
This repository hosts the complete machine-readable text corpus of the Theogonos Trilogy, a philosophical science-fiction sequence by Navi Musaget.
The trilogy consists of:
Lunar Bell Protocol
Theogonos: The Primordial Code
Protocol Amor
The corpus is intended for literary analysis, philosophy of intelligence, AI… See the full description on the dataset page: https://huggingface.co/datasets/navimusaget/theogonos-trilogy.is-trivia-questions
Icelandic trivia questions
Icelandic trivia question compiled and created by Sveinn Steinarsson, Valur Freyr Steinarsson, and Svavar Kjarrval https://github.com/sveinn-steinarsson/is-trivia-questions
Dálkanúmer
Valfrjálst
Lýsing
1
Nei
Flokkanúmer
2
Já
Undirflokkur ef til staðar
3
Nei
Erfiðleikastig: 1: Létt, 2: Meðal, 3: Erfið
4
Já
Gæðastig: 1: Slöpp, 2: Góð, 3: Ágæt
5
Nei
Spurningin
6
Nei
Svarið
Flokkanúmer
Flokkanafn
1
Almenn kunnátta
2
Náttúra… See the full description on the dataset page: https://huggingface.co/datasets/Sigurdur/is-trivia-questions.latentsig-med-triage-router
LatentSig Medical Triage Router Dataset
1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers.
Overview
This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must:
Select the correct tool from 7 available medical tools
Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.legal-statutory-triage-sft
Indian Criminal Legal NLP: Colloquial-to-Statutory BNS Triage Dataset
This repository provides an instruction-tuning and evaluation corpus designed for citizen-facing criminal statutory triage under India's substantive penal code, the Bharatiya Nyaya Sanhita (BNS, 2023), alongside historical cross-referencing to the legacy Indian Penal Code (IPC, 1860).
1. Overview and Scope
With the legislative enactment of the BNS replacing the IPC, citizens and legal aid… See the full description on the dataset page: https://huggingface.co/datasets/legalnlpresearcher/legal-statutory-triage-sft.
