bbb
intern-s1-mini-context-conditioned-molecule-transfer-v10-3-bbb-martins-bestintern-s1-mini-context-conditioned-molecule-transfer-v10-4-bbb-martins-bestintern-s1-mini-context-conditioned-molecule-transfer-v10-3-tdc-bbb-martins-bestSMILY-APE-BBBPintern-s1-mini-assay-transfer-record-level-v27-bbb-martins-l2-step1575intern-s1-mini-assay-transfer-record-level-v27-bbb-martins-l3-bestintern-s1-mini-assay-transfer-record-level-v27-bbb-martins-l4-retry1-r9nkhtrpDeepSeek-R1-Distill-Qwen-7B
Datasets
All datasets matching “bbb”StreamDelta
Streamo
Streaming Video Instruction Tuning
A real-time streaming video LLM that serves as a general-purpose interactive assistant.
📑 Paper | 🌐 Web | 🤗 Huggingface
This is the official implementation of the paper 'Streaming Video Instruction Tuning'.
News📰
[2026/2/27]:🎉Our paper has been accepted by CVPR 2026!
[2026/1/27]:🔥We have released the Streamo-Instruct dataset.[HF].
[2026/1/22]:🔥We have released our training… See the full description on the dataset page: https://huggingface.co/datasets/BBBBCHAN/StreamDelta.MoleculeNet_BBBP
MoleculeNet BBBP
BBBP (Blood-Brain Barrier Penetration) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict blood-brain barrier penetration (barrier permeability) of small drug-like molecules.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
2039
Recommended split
scaffold
Recommended metric
AUROC
References
[1]
Ines Filipa Martins et al.
"A… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BBBP.bbbnuke-screening-1B
BBB-Nuke: 1 Billion Compound Blood-Brain Barrier Permeability Screen
Overview
1.02 billion small molecules screened for blood-brain barrier (BBB) permeability using the BBB-Nuke pipeline (v0.12.0).
Screening Pipeline
Each compound was scored through:
Standardization - SMILES canonicalization via RDKit
Physicochemical properties - MW, LogP, TPSA, HBD, HBA, Fsp3, heavy atom count
pKa prediction - Representative pKa via MolGpKa (batched GCN inference)… See the full description on the dataset page: https://huggingface.co/datasets/ATTN-Lab/bbbnuke-screening-1B.BBBTTTlibero_mix_rldsTC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.
