trillion
Datasets
All datasets matching “trillion”TheBioCollection
TheBioCollection
TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.trilliongame
Bangumi Image Base of Trillion Game
This is the image base of bangumi Trillion Game, we detected 100 characters, 11831 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/trilliongame.rBridge
🌉 rBridge Paper's Reasoning Traces & Token Logprobs
This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks,
released as part of the rBridge project
(paper).
rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood
over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B)
can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.sim-scholar-eval-artifacts
eval/ — output directory index
Auto-generated map. Metrics: <dir>/<surface>/<bench>/pf_<model>.json; trajectories: <dir>/traj/.
dir
run/ckpt
benches
surfaces
#models
#valid
notes
bandps_v2_local
?
internal_citation_holdout,internal_known_item,litsearch
untagged
1
9
SEPARATE: band+paper_set v2 internal
bandps_v2_sweep
?
litqa2_validation,paper_finder_litqa2_validation,paper_finder_validation
untagged
1
9
SEPARATE: band+paper_set v2 sweep
concat_v1_local
?… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/sim-scholar-eval-artifacts.TheBioCollection-Eval
TheBioCollection-Eval
TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets.
Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.sim-scholar-qa-archive
