CoolFace
20 results

trillion

trillionlabs /TheBioCollection TheBioCollection TheBioCollection is a 52.6B-token pretraining-scale corpus for biology that transforms heterogeneous biological resources into LLM training-friendly data. It is built through a construction pipeline that collects resources across biological domains, refines them through deduplication, entity tagging and augmentation, enriches them with tool-computed biological properties, and render them as instruction-form data with programmatically verifiable answers. The… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection.texttext-generation10M<n<100M20 likes2.7k downloads2mo agoHugging FaceBangumiBase /trilliongame Bangumi Image Base of Trillion Game This is the image base of bangumi Trillion Game, we detected 100 characters, 11831 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/trilliongame.image10K<n<100K0 likes1.5k downloads1y agoHugging Facetrillionlabs /rBridge 🌉 rBridge Paper's Reasoning Traces & Token Logprobs This dataset contains GPT-4o reasoning traces and token-level logprobs for six reasoning benchmarks, released as part of the rBridge project (paper). rBridge uses these traces as gold-label reasoning references. By computing a weighted negative log-likelihood over these traces — where each token is weighted by the frontier model's confidence — small proxy models (≤1B) can reliably predict the reasoning performance of much larger… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/rBridge.tabulartext-generation10K<n<100K1 likes351 downloads7mo agoHugging Facetrillionlabs /sim-scholar-eval-artifacts eval/ — output directory index Auto-generated map. Metrics: <dir>/<surface>/<bench>/pf_<model>.json; trajectories: <dir>/traj/. dir run/ckpt benches surfaces #models #valid notes bandps_v2_local ? internal_citation_holdout,internal_known_item,litsearch untagged 1 9 SEPARATE: band+paper_set v2 internal bandps_v2_sweep ? litqa2_validation,paper_finder_litqa2_validation,paper_finder_validation untagged 1 9 SEPARATE: band+paper_set v2 sweep concat_v1_local ?… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/sim-scholar-eval-artifacts.0 likes337 downloads3mo agoHugging Facetrillionlabs /TheBioCollection-Eval TheBioCollection-Eval TheBioCollection-Eval is a biological evaluation suite for assessing large language models (BioLMs) for biology across small molecules, proteins, genomic sequences, cells/pathways, and cross-domain reasoning. It is constructed by drawing subtasks from many scattered existing benchmarks (Mol-Instructions, MolLangBench, BioReason-Pro, PerturBench Replogle K562) and combining them with source-derived newly-constructed instruction datasets. Evaluation… See the full description on the dataset page: https://huggingface.co/datasets/trillionlabs/TheBioCollection-Eval.texttext-generation1K<n<10K2 likes227 downloads3mo agoHugging Facetrillionlabs /sim-scholar-qa-archive0 likes115 downloads3mo agoHugging Face