CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01trillionlabs /TheBioCollection TheBioCollection TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.texttext-generation10M<n<100M20 likes2.7k downloads2mo agoHugging Face02trillionlabs /rBridge 🌉 rBridge Paper's Reasoning Traces & Token Logprobs This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks, released as part of the rBridge project (paper). rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B) can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.tabulartext-generation10K<n<100K1 likes351 downloads7mo agoHugging Face03trillionlabs /TheBioCollection-Eval TheBioCollection-Eval TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets. Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.texttext-generation1K<n<10K2 likes227 downloads3mo agoHugging Face04trillionlabs /SimScholar-SFT S3 SFT Trajectories Complete ReAct trajectories for scientific-literature search. Code · S3 collection · Source corpus This dataset contains 14,633 single- and two-hop tool-use trajectories. In each trajectory, a policy searches and reads a fixed scientific corpus through nine tools, then submits an answer with a correctness label. The messages column uses OpenAI tool-calling chat format. At a glance Question type Rows Correct Incorrect Single-hop… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-SFT.tabularquestion-answering10K<n<100K0 likes87 downloads2mo agoHugging Face05trillionlabs /MWMBenchtext1K<n<10K0 likes81 downloads6mo agoHugging Face06trillionlabs /NemoSlides-DPO-mix-v1.0 Slide-DPO Direct Preference Optimization dataset for training LLMs to generate slide presentations in Slidev markdown format, derived from the Slides-Align human preference rankings over the SlidesGen-Bench benchmark. Each row is a preference pair: a brief plus an available image pool as the prompt, and two Slidev-markdown responses (with <think> reasoning traces) that were generated by differently-ranked AI slide-generation products for the same brief. Row schema… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/NemoSlides-DPO-mix-v1.0.texttext-generation1K<n<10K6 likes69 downloads5mo agoHugging Face07trillionlabs /rbridge-mask rBridge-Mask Evaluation dataset for rBridge, a method for predicting LLM reasoning performance using small proxy models. Contains reasoning traces from frontier models with <span> tags marking key reasoning steps. Overview Each sample contains a question and a reasoning trace where important factual/reasoning content is tagged with <span>...</span>. rBridge computes the masked log-likelihood — only scoring tokens inside tagged regions — to predict downstream… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rbridge-mask.texttext-generation1K<n<10K3 likes60 downloads8mo agoHugging Face08trillionlabs /SimScholar-RL S3 RL Dataset Synthetic literature-search questions for agentic reinforcement learning. Code · S3 collection · Source corpus The S3 RL dataset contains 13,000 synthetic, single-hop questions for training and analyzing scientific-literature search agents. Each question is grounded in one paper from a fixed corpus of approximately 1.12 million computer-science papers. At a glance Property Value Rows 13,000 Language English Task Tool-assisted… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/SimScholar-RL.textquestion-answering10K<n<100K0 likes33 downloads2mo agoHugging Face09trillionlabs /NemoSlides-SFT-mix-v1.0 NemoSlides-SFT-mix-v1.0 NemoSlides-SFT-mix-v1.0 is a SFT training set designed to train large language models to use Slidev to generate slides. text1K<n<10K5 likes32 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.