datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
smiles-transformers
smiles-transformers dataset
TODO: Add references to the datasets we curated
dataset features
name: text
Molecule SMILES : string
name: formula
Molecular formula : string
name: NumHDonors
Number of hidrogen bond donors : int
name: NumHAcceptors
Number of hidrogen bond acceptors : int
name: MolLogP
Wildman-Crippen LogP : float
name: NumHeteroatoms
Number of hetero atoms: int
name: RingCount
Number of rings : int
name: NumRotatableBonds
Number of rotable… See the full description on the dataset page: https://huggingface.co/datasets/maykcaldas/smiles-transformers.kor_unsmileSWE-Smith-Seeds-Clean
SWE-Smith Seeds, agent-verified
1,552 of SWE-smith's 59,136 instances, repackaged as terminal tasks and kept only where every claim about them was executed and held: the bug is present, the reference fix earns the grader's reward, the repository's own suite still passes, and a coding agent solved the task from its instruction alone in a sandbox that had neither the fix nor the tests nor the network. Every row carries the verdict and the conditions it was taken under; nothing… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/SWE-Smith-Seeds-Clean.smiles_transformer_shuffledswe_smith_js_qwen3.5_35b_trajs_4358swe_smith_rebenchv2_5136
SWE-smith + SWE-rebench V2 5136 Mix
This dataset is the swe_smith_rebenchv2_5136 training mix used by the rLLM SWE training scripts. It combines filtered SWE-smith trajectory tasks with sampled SWE-rebench V2 tasks so future training jobs can pull the prepared parquet directly instead of regenerating it with the long preparation script.
Contents
data/train.parquet: the canonical rLLM task rows, 5,136 examples.
rllm_verl/train.parquet: the rLLM DatasetRegistry… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/swe_smith_rebenchv2_5136.swe_smith_java_qwen3.5_35b_trajs_4369galaxies_metadata
Galaxy metadata for pairing with `smith42/galaxies' dataset
Here we have metadata for ~8.5 million galaxies.
This metadata can be paired with galaxy jpg cutouts from the DESI legacy survey DR8,
the cut outs are found here: https://huggingface.co/datasets/Smith42/galaxies.
I've split away 1% of the metadata into a test set, and 1% into a validation set.
The remaining 98% of the metadata comprise the training set.
Useful links
Paper here:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxies_metadata.smints2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 51,
"total_frames": 13210,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Unseq/smints2.chembl-smiles-curated
Dataset Card for Curated ChEMBL SMILES Dataset (lovingscience/chembl-smiles-curated)
This dataset is a reproducibly cleaned, deduplicated, and split ChEMBL SMILES dataset built for training molecular generative models and chemical space mapping tools (such as Imolgen and Navifold).
Dataset Pipeline Provenance & Configuration
ChEMBL Release Version: 37
RDKit Version: 2026.03.6
Deduplication & Split Seed: 42
Split Ratios: Train (80.0%) / Validation (10.0%) / Test… See the full description on the dataset page: https://huggingface.co/datasets/lovingscience/chembl-smiles-curated.so-100-draw-smileyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 53,
"total_frames": 24614,
"total_tasks": 1,
"total_videos": 106,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:53"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kevin510/so-100-draw-smiley.SWE-smith-oracle-4k-context-1k-diffgroot1_smileface_Prod_try2
groot1_smileface_Prod_try2
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
chembl-v37-smiles-curated
Dataset Card for Curated ChEMBL SMILES Dataset (lovingscience/chembl-smiles-curated)
This dataset is a reproducibly cleaned, deduplicated, and split ChEMBL SMILES dataset built for training molecular generative models and chemical space mapping tools (such as Imolgen and Navifold).
Dataset Pipeline Provenance & Configuration
ChEMBL Release Version: 37
RDKit Version: 2026.03.6
Deduplication & Split Seed: 42
Split Ratios: Train (80.0%) / Validation (10.0%) / Test… See the full description on the dataset page: https://huggingface.co/datasets/lovingscience/chembl-v37-smiles-curated.chembl-2025-randomized-smiles-cleaned-rdkit-descriptorsswe_smith_py_qwen3.5_35b_trajs_3934_no_p2pafrica-smishing-sms-phishing
SMS Phishing / Smishing (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-smishing-sms-phishing.desi_hsc_crossmatched
Crossmatched samples from the Multimodal Universe for DESI/HSC
Mother paper here: https://arxiv.org/abs/2412.02527
galaxies_with_embeddingshalf-of-chembl-2025-randomized-smiles-cleaned-rdkit-descriptorsR2E-Smithswe_smith_rs_qwen3.5_35b_trajs_2477sdss_gaia_crossmatched
Crossmatched samples from the Multimodal Universe for SDSS/Gaia
Mother paper here: https://arxiv.org/abs/2412.02527
swe_smith_py_qwen3.5_35b_trajs_1952swe_smith_go_qwen3.5_35b_trajs_1448swe_smith_go_qwen3.5_35b_trajs_1629_no_p2pdesi_sdss_crossmatched
Crossmatched samples from the Multimodal Universe for DESI/SDSS
Mother paper here: https://arxiv.org/abs/2412.02527
Smilodonswe_smith_go_trajs_1629_incompleteSmilodonUnif
