drugs
Datasets
All datasets matching “drugs”drug-seq-u2os-novartisI AM NOT AFFILIATED WITH NOVARTIS IN ANY WAY; THIS IS SIMPLY AN UPLOAD OF THEIR DATASET, "NOVARTIS/DRUG-SEQ U2OS MOABOX DATASET."
Novartis DRUG-seq U2OS MoABox Dataset
This dataset profiles transcriptomic responses of the U-2 OS human osteosarcoma cell line to a broad collection of small molecule perturbations. It contains 49,392 observations spanning 3,742 unique compounds tested at 4 distinct dosages + 0.0, each annotated with their respective mechanisms of action (MoA).
Each… See the full description on the dataset page: https://huggingface.co/datasets/TitouanCh/drug-seq-u2os-novartis.FR_Drugs_tiret_dictation_augmented
FR_Drugs_tiret_dictation_augmented
Dictee de listes de medicaments au format "tiret" (patterns du dataset
FR_Drugs_dictation_with_dashes_pattern, molecules completees depuis les sources
autocorrect/medical), synthetisee en TTS Coqui XTTS v2 (600 voix clonees) puis
augmentee acoustiquement.
element
valeur
extraits
141960
moteur
Coqui XTTS v2, 600 voix clonees (16 kHz mono)
base
audio/coqui/<shard>/
parasite seul (65%)
audio_babble_only/
echo seul (25%)… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/FR_Drugs_tiret_dictation_augmented.geom_drugs
GEOM: Molecular Conformations (Drugs Subset)
Note: This is a mirrored and specifically preprocessed version of the GEOM dataset (Drugs subset), originally created by Simon Axelrod and Rafael Gómez-Bombarelli. All credit for the original conformational sampling and DFT calculations goes to the original authors. This repository exists to guarantee availability and exact reproducibility for downstream machine learning projects.
Dataset Description
The Geometric Ensemble… See the full description on the dataset page: https://huggingface.co/datasets/raulsofia/geom_drugs.TDC_pampa_approved_drugsMisssing_FR_TTS_ICD_Snomed_drugs_0806
Missing FR TTS ICD Snomed drugs 0506
Dataset TTS pour des termes medicaux manquants. Genere le 2026-06-06.
Contenu
53192 fichiers audio
25295 textes uniques
Edge TTS: 25277 fichiers (voix Denise, Microsoft)
Coqui XTTS: 0 fichiers (2 voix/texte, ~10% des textes)
Source: synthese sur 3x V100-32GB
Format
Les fichiers audio sont dans archives/*.tar.gz.
Le fichier metadata.jsonl est a la racine du dataset.
Stats
Edge TTS: 25277
Coqui… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/Misssing_FR_TTS_ICD_Snomed_drugs_0806.GEOM-DRUGS_ADiT
All-atom Diffusion Transformers - GEOM-DRUGS dataset
GEOM-DRUGS dataset from the paper "All-atom Diffusion Transformers: Unified generative modelling of molecules and materials", by Chaitanya K. Joshi, Xiang Fu, Yi-Lun Liao, Vahe Gharakhanyan, Benjamin Kurt Miller, Anuroop Sriram*, and Zachary W. Ulissi* from FAIR Chemistry at Meta (* Joint last author).
Original data source:
https://github.com/cvignac/MiDi?tab=readme-ov-file#datasets
https://github.com/learningmatter-mit/geom… See the full description on the dataset page: https://huggingface.co/datasets/chaitjo/GEOM-DRUGS_ADiT.
