datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ais-hudson-rivercute-kitchen-d8edf8
cute-kitchen-d8edf8
Synthetic products test data: 49 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hudsonmichael235/cute-kitchen-d8edf8.Resume-Datasethudson-2023-dosedo
hudson-2023-dosedo
SMILES of ~3.7 million diversity-oriented synthesis (DOS) compounds from a reported DNA-encoded library, in:
Hudson, L., Mason, J.W., Westphal, M.V. et al. Diversity-oriented synthesis encoded by deoxyoligonucleotides.
Nat Commun 14, 4930 (2023). https://doi.org/10.1038/s41467-023-40575-5
The SMILES strings have been canonicalized, and split into training (70%), validation (15%), and test (15%) sets by Murcko scaffold. Additional features like molecular weight… See the full description on the dataset page: https://huggingface.co/datasets/scbirlab/hudson-2023-dosedo.newdatanews.csvfnews
