SBD
Datasets
All datasets matching “SBD”sbd-qr-subset
SBD QR Subset — low resolution
A mirror of the low-resolution ROI split of the Synthetic Barcode Dataset
(Quenum, Wang, Zakhor), repackaged from 749,682 loose files into parquet.
split
ROIs
instances
train
80,000
439,731
validation
10,000
55,072
test
10,000
54,876
total
100,000
549,679
Why this repackaging exists
Upstream, this split is three-quarters of a million individual PNG and JPEG
files. That is unpleasant to move, impossible to browse… See the full description on the dataset page: https://huggingface.co/datasets/devmandan/sbd-qr-subset.UPRPRC_SBD_KVrna-sbdd-v2
RNA-SBDD v2
A frozen benchmark for RNA structure-based drug design: 8,006 RNA pocket-ligand
complexes derived from RCSB, with a sequence-identity-disjoint split, the
evaluation artifacts, and the trained checkpoints the benchmark's numbers come
from.
This repository exists because the cluster the work ran on was retired. It is a
complete handoff — dataset, artifacts, weights, and the tooling to bring all of
it up somewhere else.
Code: git@Ced3-han:Ced3-han/RNASBDD.git, branch… See the full description on the dataset page: https://huggingface.co/datasets/CedLJH/rna-sbdd-v2.chagatai-sbd
Chagatai Sentence Boundary Detection
Canonical word-level Sentence Boundary Detection data for Chagatai. South
Uzbek (uzs) and Uyghur (uig) are optional train-only auxiliary languages.
Every configuration uses the same Chagatai train source split. Validation and
test are physically shared files referenced by all five configurations.
Load with datasets
from datasets import load_dataset
dataset = load_dataset("chagatai-project/chagatai-sbd", "chagatai_only")… See the full description on the dataset page: https://huggingface.co/datasets/chagatai-project/chagatai-sbd.Practical_SBDD
PDBBind.lmdb.zip
processed pdbbind data for training in lmdb format. Docs for lmdb can be found at: https://lmdb.readthedocs.io/en/release/
PDBBind-DUD_E_FLAPP_0.6.pkl
train/valid split file for 0.6 version
PDBBind-DUD_E_FLAPP_0.9.pkl
train/valid split file for 0.9 version
DUDE.zip
DUD-E test set. Each directory is a target and contains all needed files for evaluation.
LIT-PCBA.zip
LIT-PCBA test set. Each directory is a target and… See the full description on the dataset page: https://huggingface.co/datasets/bgao95/Practical_SBDD.Synth-SBDH
Dataset Card for Synth-SBDH
Synth-SBDH is a collection of 8,767 synthetic examples with annotations for 15 SBDH categories. SBDH annotations include information such as presence, period and annotation rationale.
Dataset Description
Synth-SBDH is a novel synthetic SBDH dataset that mimics EHR notes.
Repository: Codes to reproduce experiments
Paper: Link
Point of Contact: Avijit Mitra
Dataset Structure
Data Instances
Some examples from… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/Synth-SBDH.
