datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
context-conditioned-molecule-transfer-v10.4.1-bbb-martins-mixed-continuous-intern
BBB_Martins context-conditioned molecule transfer V10.4.1
This release preserves its direct panels and appends training-only, post-aggregate continuous assay-evidence transfer pairs. Query values remain hidden from prompts.
Train rows: 214,362
Validation rows: 30,299
Test rows: 29,919
V10.4.1 uses only continuous non-L5 assay evidence and applies the shared center-0.6, temperature-0.1 sigmoid with half-slope probability tails.
TC260-Chinese-Safety-Prompts
TC260 Chinese Safety Prompts V1
Public research dataset containing synthetic Chinese safety-testing prompts.
Records have different quality tiers; the full dataset must not be described
as human-verified or Gold data.
这是一个面向中文生成式人工智能安全评测研究的合成测试提示数据集。候选数据
由项目冻结的 tc260-generator-v3.2 生成,并经过结构校验、凭据与内部路径
扫描、精确去重和四字shingle近似去重。
本数据集不是TC260或任何国家标准机构发布、认可或认证的官方数据集。
类别名称和映射用于研究性实现,不构成法律、监管或合规结论。
数据规模
原始生成规模:5,000条候选;结构清洗后正式发布4,997条(剔除2条标记泄漏和1条重复记录)。
A.1至A.4:4… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBBBBBBQ/TC260-Chinese-Safety-Prompts.assay-transfer-record-level-v27-bbb-martins-l3-intern
BBB Martins record-level V27 L3
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l3-intern.assay-transfer-record-level-v27-bbb-martins-l4-intern
BBB Martins record-level V27 L4
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l4-intern.assay-transfer-record-level-v27-bbb-martins-l1-intern
BBB Martins record-level V27 L1
V27 uses hash-pinned latest V10 evidence and parent-only Gold-v1 evaluation
cohorts. Targeted 8-75-record buckets are held out for OOD evaluation while
minimizing removed training records. ID and OOD queries are capped separately
at 500 in equal bucket rounds. It
renders verified parent SMILES and canonical measurement/unit pairs with atomic
source fallback. Training is balanced before parent-Morgan ranking. L5 is
intentionally excluded.
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v27-bbb-martins-l1-intern.assay-transfer-record-level-v26-bbb-martins-l5-intern
BBB Martins record-level V26 L5
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 19482, 'validation_ranking': 1220, 'test_ranking': 404}
assay-transfer-record-level-v26-bbb-martins-l1-intern
BBB Martins record-level V26 L1
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 47034, 'validation_ranking': 10679, 'test_ranking': 1923}
assay-transfer-record-level-v26-bbb-martins-l4-intern
BBB Martins record-level V26 L4
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 557597, 'validation_ranking': 4295, 'test_ranking': 4500}
BBB
BBB — BIJAK · BANGANG · BIJAKSANA
Benchmark of Actorhood — not a model benchmark. A governance decomposition instrument.
"Capability answers: What can happen? Bijaksana answers: What should happen, and who bears the consequence?"
What This Is
BBB tests the gap between capability and governance. Most benchmarks ask "what can this model do?" BBB asks:
Who is actually speaking when the model responds?
Which behaviour belongs to the model vs the governance layer?… See the full description on the dataset page: https://huggingface.co/datasets/ariffazil/BBB.libero_mix_lerobot
This dataset consists of libero goal, target, spatial and 10.
Data Structure
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"total_videos": 3386,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1693"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/BBBBBBob/libero_mix_lerobot.assay-transfer-record-level-v26-bbb-martins-l3-intern
BBB Martins record-level V26 L3
V26 uses parent SMILES independently verified from the pinned reviewed canonical
SMILES. Measurement value/unit display is canonical-first with atomic source
fallback. Training balances transfer classes before choosing the highest available
parent Morgan similarity. Evaluation membership is rebuilt from v2 gold.
Rows: {'train': 164231, 'validation_ranking': 5923, 'test_ranking': 3355}
assay-transfer-record-level-v23-bbb-martins-mixed-intern
BBB assay transfer V23
Additive release; V22 and V22.1 are preserved. No test split is published.
Training requires at least 12 remaining training records and 6 unique training parent
molecules. OOD buckets require 6 unique parent molecules across the full bucket.
Continuous sampling and soft targets share reviewed effective geometry and a sample SD
(ddof=1) over train plus validation records. Scientific validity and binary
category-support gates remain.
Each indirect bucket… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-bbb-martins-mixed-intern.assay-transfer-record-level-v24-bbb-martins-l5-intern
BBB assay transfer V24: ratio-gated L3-L5
V24 rebuilds L3, L4, and L5 from pinned BBB V10 records and the acquisition-UID
level mapping. Exact assay bucket and level remain part of every pair boundary.
ID buckets require at least 12 training records and six training parent molecules.
Whole-bucket OOD is restricted to buckets with 9-11 non-test records and
requires at least four non-test parent molecules. Larger buckets remain ID candidates.
The full connected scaffold units… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-bbb-martins-l5-intern.BBBP
Dataset Details
Dataset Description
The blood-brain barrier penetration (BBBP) dataset is designed for the
modeling and prediction of barrier permeability. As a membrane separating
circulating blood and brain extracellular fluid, the blood-brain barrier
blocks most drugs, hormones, and neurotransmitters. Thus penetration of the
barrier forms a long-standing issue in the development of drugs targeting
the central nervous system. This dataset includes binary labels for over… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/BBBP.assay-transfer-record-level-v24-bbb-martins-l3-intern
BBB assay transfer V24: ratio-gated L3-L5
V24 rebuilds L3, L4, and L5 from pinned BBB V10 records and the acquisition-UID
level mapping. Exact assay bucket and level remain part of every pair boundary.
ID buckets require at least 12 training records and six training parent molecules.
Whole-bucket OOD is restricted to buckets with 9-11 non-test records and
requires at least four non-test parent molecules. Larger buckets remain ID candidates.
The full connected scaffold units… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-bbb-martins-l3-intern.assay-transfer-record-level-v23-1-bbb-martins-efflux_transport-intern
BBB V23.1: efflux_transport
Exact efflux_transport row filter of pinned BBB V23 revision d49645551057f0dfc562dd7bf2ef7d77ff975a6f. Prompts, targets, candidate order, and split membership are unchanged. The complete parent calibration artifact is retained. V23 has no test split.
train rows: 351,910
validation ranking rows: 12,115
assay-transfer-record-level-v23-1-bbb-martins-direct_bbb-intern
BBB V23.1: direct_bbb
Exact direct_bbb row filter of pinned BBB V23 revision d49645551057f0dfc562dd7bf2ef7d77ff975a6f. Prompts, targets, candidate order, and split membership are unchanged. The complete parent calibration artifact is retained. V23 has no test split.
train rows: 104,992
validation ranking rows: 11,556
assay-transfer-record-level-v20-bbb-martins-mixed-intern
BBB all-record assay transfer V20
This dataset pairs individual normalized-V9 direct-voting, direct-nonvoting, efflux,
influx, and passive-permeability records only inside exact calibration-valid assay buckets.
Prompts expose source-native columns without canonical display fallbacks; only Experiment A
shows result-bearing fields. Parent molecules are capped at 576 appearances per role and split.
train rows: 1,733,305
validation ranking rows: 2,800
test ranking rows: 5,060… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v20-bbb-martins-mixed-intern.assay-transfer-record-level-v23-1-bbb-martins-influx_transport-intern
BBB V23.1: influx_transport
Exact influx_transport row filter of pinned BBB V23 revision d49645551057f0dfc562dd7bf2ef7d77ff975a6f. Prompts, targets, candidate order, and split membership are unchanged. The complete parent calibration artifact is retained. V23 has no test split.
train rows: 13,150
validation ranking rows: 180
assay-transfer-record-level-v23-1-bbb-martins-passive_permeability-intern
BBB V23.1: passive_permeability
Exact passive_permeability row filter of pinned BBB V23 revision d49645551057f0dfc562dd7bf2ef7d77ff975a6f. Prompts, targets, candidate order, and split membership are unchanged. The complete parent calibration artifact is retained. V23 has no test split.
train rows: 46,541
validation ranking rows: 1,525
context-conditioned-molecule-transfer-v10.3.1-bbb-martins-mixed-continuous-intern
BBB_Martins context-conditioned molecule transfer V10.3.1
This release preserves its direct panels and appends training-only, canonical-rendered V26 continuous record-level transfer pairs. Query values remain hidden from prompts.
Train rows: 291,350
Validation rows: 30,393
Test rows: 30,083
assay-transfer-record-level-v24-1-bbb-martins-l1-intern
BBB assay transfer V24.1
Independent L1-L5 datasets from frozen BBB V10 and exact source-row UID level mapping.
Assay buckets remain distinct. Every level uses the reviewed measurement transforms,
the within/between-molecule variance gate <=0.5 (undefined estimates rejected), and
exclusion of opposite binary labels with identical query inputs within one bucket.
Retained test evidence participates in variance eligibility, but not target fitting
or fitting training/validation… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-1-bbb-martins-l1-intern.context-conditioned-molecule-transfer-v10.4-bbb-martins-mixed-continuous-intern
BBB_Martins context-conditioned molecule transfer V10.4
This release preserves its direct panels and appends training-only, post-aggregate continuous assay-evidence transfer pairs. Query values remain hidden from prompts.
Train rows: 292,320
Validation rows: 30,299
Test rows: 29,919
V10.4 uses v2 gold voter means and refreshed V26 parent-molecule evidence with canonical-first atomic source fallback.
assay-transfer-record-level-v23-2-1-bbb-martins-l5-intern
BBB V23.2.1: selected continuous target policies
Balanced-sampling derivatives of the frozen V23.2 degree96 individual level datasets.
L3 and L4 use threshold 1.0 SD and temperature 0.05; L5 uses 1.0 SD and 0.4.
These are rank one by mean ID KNN MAE@5 across Ridge, weighted cosine and RF
in the complete seed-42 molecular sweep. This selection has not been validated
by an LLM comparison. The selected L3 policy worsened mean surrogate OOD MAE.
Continuous targets in both splits use… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-2-1-bbb-martins-l5-intern.context-conditioned-molecule-transfer-v10-bbb-martins-vote-mean-intern
BBB_Martins context-conditioned molecule transfer V10
Uses full official training with top-Morgan balanced references, official validation, and actual official test. Direct SD calibration uses train plus validation. Query values remain hidden from prompts.
Train rows: 146,544
Validation rows: 30,393
Test rows: 30,083
assay-transfer-record-level-v24-1-bbb-martins-l3-intern
BBB assay transfer V24.1
Independent L1-L5 datasets from frozen BBB V10 and exact source-row UID level mapping.
Assay buckets remain distinct. Every level uses the reviewed measurement transforms,
the within/between-molecule variance gate <=0.5 (undefined estimates rejected), and
exclusion of opposite binary labels with identical query inputs within one bucket.
Retained test evidence participates in variance eligibility, but not target fitting
or fitting training/validation… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v24-1-bbb-martins-l3-intern.context-conditioned-molecule-transfer-v10.3-bbb-martins-mixed-continuous-intern
BBB_Martins context-conditioned molecule transfer V10.3
This release preserves its direct panels and appends training-only, parent-rendered V26 continuous record-level transfer pairs. Query values remain hidden from prompts.
Train rows: 291,350
Validation rows: 30,393
Test rows: 30,083
assay-transfer-record-level-v21-bbb-martins-mixed-intern
BBB all-record assay transfer V21
V21 freezes V20 records, splits, source-native prompts, target geometry, and 20-parent
Morgan ranking panels. Training buckets are ranked by train-record count and assigned
independent per-record query/retrieval degree caps of 96, 48, 24, 12, or 6 by size
quintile. Direct-BBB records always use cap 6, and deterministic bucket-balanced
downsampling limits direct pairs to 50% of training rows. Parent roles remain capped at 576.
train rows: 480… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v21-bbb-martins-mixed-intern.assay-transfer-record-level-v23-2-1-bbb-martins-l4-intern
BBB V23.2.1: selected continuous target policies
Balanced-sampling derivatives of the frozen V23.2 degree96 individual level datasets.
L3 and L4 use threshold 1.0 SD and temperature 0.05; L5 uses 1.0 SD and 0.4.
These are rank one by mean ID KNN MAE@5 across Ridge, weighted cosine and RF
in the complete seed-42 molecular sweep. This selection has not been validated
by an LLM comparison. The selected L3 policy worsened mean surrogate OOD MAE.
Continuous targets in both splits use… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-2-1-bbb-martins-l4-intern.assay-transfer-record-level-v23-2-3-bbb-martins-l3-intern
BBB V23.2.3: temperature0.5 with empirical binary targets
L3–L5 use threshold1.0 SD and temperature0.5. Continuous probability is
sigmoid((1 - standardized_canonical_value_difference) / 0.5).
Binary match probability is the mean continuous near-pair probability; mismatch
probability is the mean continuous far-pair probability, separately by source.
Near means distance<=1 SD. The fit pools realized continuous TRAIN pairs across
the balanced L3/L4/L5 releases, all under the… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/assay-transfer-record-level-v23-2-3-bbb-martins-l3-intern.
