datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreamDelta
Streamo
Streaming Video Instruction Tuning
A real-time streaming video LLM that serves as a general-purpose interactive assistant.
📑 Paper | 🌐 Web | 🤗 Huggingface
This is the official implementation of the paper 'Streaming Video Instruction Tuning'.
News📰
[2026/2/27]:🎉Our paper has been accepted by CVPR 2026!
[2026/1/27]:🔥We have released the Streamo-Instruct dataset.[HF].
[2026/1/22]:🔥We have released our training… See the full description on the dataset page: https://huggingface.co/datasets/BBBBCHAN/StreamDelta.MoleculeNet_BBBP
MoleculeNet BBBP
BBBP (Blood-Brain Barrier Penetration) dataset [1], part of MoleculeNet [2] benchmark. It is intended to be used through
scikit-fingerprints library.
The task is to predict blood-brain barrier penetration (barrier permeability) of small drug-like molecules.
Characteristic
Description
Tasks
1
Task type
classification
Total samples
2039
Recommended split
scaffold
Recommended metric
AUROC
References
[1]
Ines Filipa Martins et al.
"A… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BBBP.bbbnuke-screening-1B
BBB-Nuke: 1 Billion Compound Blood-Brain Barrier Permeability Screen
Overview
1.02 billion small molecules screened for blood-brain barrier (BBB) permeability using the BBB-Nuke pipeline (v0.12.0).
Screening Pipeline
Each compound was scored through:
Standardization - SMILES canonicalization via RDKit
Physicochemical properties - MW, LogP, TPSA, HBD, HBA, Fsp3, heavy atom count
pKa prediction - Representative pKa via MolGpKa (batched GCN inference)… See the full description on the dataset page: https://huggingface.co/datasets/ATTN-Lab/bbbnuke-screening-1B.BBBTTTlibero_mix_rldscontext-conditioned-molecule-transfer-v10.4.1-bbb-martins-mixed-continuous-intern
BBB_Martins context-conditioned molecule transfer V10.4.1
This release preserves its direct panels and appends training-only, post-aggregate continuous assay-evidence transfer pairs. Query values remain hidden from prompts.
Train rows: 214,362
Validation rows: 30,299
Test rows: 29,919
V10.4.1 uses only continuous non-L5 assay evidence and applies the shared center-0.6, temperature-0.1 sigmoid with half-slope probability tails.
TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.assay-transfer-record-level-v27-bbb-martins-l3-intern
BBB Martins record-level V27 L3
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l3-intern.assay-transfer-record-level-v27-bbb-martins-l2-intern
BBB Martins record-level V27 L2
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l2-intern.assay-transfer-record-level-v27-bbb-martins-l4-intern
BBB Martins record-level V27 L4
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l4-intern.assay-transfer-record-level-v27-bbb-martins-l1-intern
BBB Martins record-level V27 L1
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l1-intern.digit-force-estimation
Dataset Details
This dataset contains paired tactile and force data, intended for use in predicting 3-axis normal and shear forces applied to the sensor's elastomer. We used three different indenter shapes to collect force-labeled data: hemisphere, sharp, and flat. To measure force ground truths, we employed the ATI nano17 force/torque sensor. The protocol consisted of applying a random normal load (up to 5N) followed by a shear load, achieved by sliding the probe 2mm on the… See the full description on the dataset page: https://huggingface.co/datasets/Bbbbbhyyyy/digit-force-estimation.AIR
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/bbbear578/AIR.assay-transfer-record-level-v26-bbb-martins-l5-intern
BBB Martins record-level V26 L5
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 19482, 'validation_ranking': 1220, 'test_ranking': 404}
assay-transfer-record-level-v26-bbb-martins-l1-intern
BBB Martins record-level V26 L1
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 47034, 'validation_ranking': 10679, 'test_ranking': 1923}
assay-transfer-record-level-v26-bbb-martins-l4-intern
BBB Martins record-level V26 L4
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 557597, 'validation_ranking': 4295, 'test_ranking': 4500}
BBB
BBB — BIJAK · BANGANG · BIJAKSANA
Benchmark of Actorhood — not a model benchmark. A governance decomposition instrument.
"Capability answers: What can happen? Bijaksana answers: What should happen, and who bears the consequence?"
What This Is
BBB tests the gap between capability and governance. Most benchmarks ask "what can this model do?" BBB asks:
Who is actually speaking when the model responds?
Which behaviour belongs to the model vs the governance layer?… See the full description on the dataset page: https://huggingface.co/datasets/ariffazil/BBB.libero_mix_lerobot
This dataset consists of libero goal, target, spatial and 10.
Data Structure
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"total_videos": 3386,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1693"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBob/libero_mix_lerobot.assay-transfer-record-level-v26-bbb-martins-l2-intern
BBB Martins record-level V26 L2
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 2187633, 'validation_ranking': 103788, 'test_ranking': 2700}
assay-transfer-record-level-v26-bbb-martins-l3-intern
BBB Martins record-level V26 L3
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 164231, 'validation_ranking': 5923, 'test_ranking': 3355}
bbbgassay-transfer-record-level-v23-bbb-martins-mixed-intern
BBB assay transfer V23
Additive release; V22 and V22.1 are preserved. No test split is published.
Training requires at least 12 remaining training records and 6 unique training parent
molecules. OOD buckets require 6 unique parent molecules across the full bucket.
Continuous sampling and soft targets share reviewed effective geometry and a sample SD
(ddof=1) over train plus validation records. Scientific validity and binary
category-support gates remain.
Each indirect bucket… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-bbb-martins-mixed-intern.NL-ReferNL-Refer Dataset
A Natural Language Referring Dataset for Fine-grained Video Object Understanding
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Overview
NL-Refer is a video object-level instruction dataset built on top of VideoRefer-700K. While the original VideoRefer uses visual prompts (colored masks overlaid on video frames) to indicate target objects, NL-Refer replaces them with natural language… See the full description on the dataset page: https://huggingface.co/datasets/BBBBCHAN/NL-Refer.bbbp
Dataset Card for bbbp
Dataset Summary
bbbp is a dataset included in MoleculeNet. This dataset has binary labels of blood-brain barrier penetration(permeability).
Dataset Structure
Data Fields
Each split contains
smiles: the SMILES representation of a molecule
selfies: the SELFIES representation of a molecule
target: blood-brain barrier penetration(permeability)
Data Splits
The dataset is split into an 80/10/10 train/valid/test split… See the full description on the dataset page: https://huggingface.co/datasets/zpn/bbbp.assay-transfer-record-level-v24-bbb-martins-l5-intern
BBB assay transfer V24: ratio-gated L3-L5
V24 rebuilds L3, L4, and L5 from pinned BBB V10 records and the acquisition-UID
level mapping. Exact assay bucket and level remain part of every pair boundary.
ID buckets require at least 12 training records and six training parent molecules.
Whole-bucket OOD is restricted to buckets with 9-11 non-test records and
requires at least four non-test parent molecules. Larger buckets remain ID candidates.
The full connected scaffold units… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-bbb-martins-l5-intern.bbbc021reasoningbbbc039-cellposeBBBP
Dataset Details
Dataset Description
The blood-brain barrier penetration (BBBP) dataset is designed for the
modeling and prediction of barrier permeability. As a membrane separating
circulating blood and brain extracellular fluid, the blood-brain barrier
blocks most drugs, hormones, and neurotransmitters. Thus penetration of the
barrier forms a long-standing issue in the development of drugs targeting
the central nervous system. This dataset includes binary labels for over… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/BBBP.Data08BBBRFM
