datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TheBioCollection
TheBioCollection
TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.rBridge
🌉 rBridge Paper's Reasoning Traces & Token Logprobs
This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks,
released as part of the rBridge project
(paper).
rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood
over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B)
can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.TheBioCollection-Eval
TheBioCollection-Eval
TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets.
Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.NemoSlides-DPO-mix-v1.0
Slide-DPO
Direct Preference Optimization dataset for training LLMs to generate slide
presentations in Slidev markdown format, derived from the
Slides-Align human
preference rankings over the
SlidesGen-Bench benchmark.
Each row is a preference pair: a brief plus an available image pool as the
prompt, and two Slidev-markdown responses (with <think> reasoning traces)
that were generated by differently-ranked AI slide-generation products for
the same brief.
Row schema… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/NemoSlides-DPO-mix-v1.0.
