datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mataian-moses-stac
Note: This description was drafted with AI assistance and is subject to revision.
MOSES Initiative — Matai'an Open Science and Engineering Sharing Platform
The MOSES Initiative (Mataian Open Science and Engineering Sharing Initiative) publishes scientific and engineering datasets for the Matai'an Creek (馬太鞍溪) watershed in eastern Taiwan (23.56°N–24.15°N, 121.16°E–121.68°E). All datasets are discoverable through a STAC catalog.
Datasets
1. Radar… See the full description on the dataset page: https://huggingface.co/datasets/NTU-CompHydroMet-Lab/mataian-moses-stac.MOSESmoses
Molecular Sets (MOSES): A benchmarking platform for molecular generation models
Deep generative models are rapidly becoming popular for the discovery of new molecules and materials. Such models learn on a large collection of molecular structures and produce novel compounds. In this work, we introduce Molecular Sets (MOSES), a benchmarking platform to support research on machine learning for drug discovery. MOSES implements several popular molecular generation models and provides a… See the full description on the dataset page: https://huggingface.co/datasets/katielink/moses.Novax_krio_ASR_v0.1Novax_krio_TTS_v0.130_hours_krio_voiceKrio_data_tts_v2newonly2-binance-aggtrades-backupmoses
Dataset Details
Dataset Description
Molecular Sets (MOSES) is a benchmark platform
for distribution learning based molecule generation.
Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization.
It is processed from the ZINC Clean Leads dataset.
Curated by:
License: CC BY 4.0
Dataset Sources
Article about original dataset
Link to publication of associated dataset - zinc
Github repository concerning the dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/moses.smiles-molecules-moses
MOSES Molecule Generation Dataset
Dataset Description
Molecular Sets (MOSES) is a benchmark platform for distribution learning based molecule generation. Within this benchmark, MOSES provides a cleaned dataset of molecules that are ideal of optimization. It is processed from the ZINC Clean Leads dataset.
Task Description
For both distribution learning-based and goal-oriented molecule generation. That is to generate new molecules that has desirable properties… See the full description on the dataset page: https://huggingface.co/datasets/antoinebcx/smiles-molecules-moses.arabic-agent-eval
Arabic Agent Eval — Dataset Card
An open, installable Arabic function-calling benchmark with dialect splits.
Dataset summary
51 evaluation items spanning 6 categories and 5 dialects of Arabic, testing whether large language models can (a) select the right tool, (b) extract arguments from natural Arabic instructions, (c) preserve Arabic text in tool arguments instead of transliterating, and (d) understand dialectal framing.
Supported tasks… See the full description on the dataset page: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval.switchboard-tierb-codeswitch
SwitchBoard Tier B — African code-switched speech
87 consented utterances of intra-sentential code-switching — Nigerian Pidgin,
Yorùbá, Hausa and Kiswahili each mixed with English inside a single sentence —
recorded from 8 bilingual volunteers at the Deep Learning Indaba 2026, Lagos.
Collected for the MLC (Africa) × Intron Agentic Voice AI Challenge as an
evaluation set for telco/fintech voice agents. 8.75 minutes total.
What this is for
Measuring whether a speech… See the full description on the dataset page: https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch.Novax_Multi_speakerInstruct-dataset11Mhausa_001foundry_moses_v1-1
Molecular Sets (MOSES): A Benchmarking Platform for Molecular Generation Models
Dataset Information
Source: Foundry-ML
DOI: 10.18126/rp13-3k3h
Year: 2022
Authors: Polykovskiy, Daniil, Zhebrak, Alexander, Sanchez-Lengeling, Benjamin, Golovanov, Sergey, Tatanov, Oktai, Belyaev, Stanislav, Kurbanov, Rauf, Artamonov, Aleksey, Aladinskiy, Vladimir, Veselov, Mark, Kadurin, Artur, Johansson, Simon, Chen, Hongming, Nikolenko, Sergey, Aspuru-Guzik, Alan, Zhavoronkov, Alex
Data… See the full description on the dataset page: https://huggingface.co/datasets/foundry-ml/foundry_moses_v1-1.tau-voice-bairong-providermoses_test
nico8771/moses_test — cleaned MOSES (test split)
Each row is a molecule as canonical SMILES plus RDKit-recomputed targets. MOSES is
neutral by construction, so molecules are featurized over 7 atom types with no
formal charges, and aromatic bonds are kept as their own class (no kekulization)
so the model learns aromaticity directly.
Source: official MOSES test.csv.gz (molecularsets/moses). Code:
https://github.com/Nico-Conti/flow-matching-molecules (dataset/).… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_test.moses_train
nico8771/moses_train — cleaned MOSES (train split)
Each row is a molecule as canonical SMILES plus RDKit-recomputed targets. MOSES is
neutral by construction, so molecules are featurized over 7 atom types with no
formal charges, and aromatic bonds are kept as their own class (no kekulization)
so the model learns aromaticity directly.
Source: official MOSES train.csv.gz (molecularsets/moses). Code:
https://github.com/Nico-Conti/flow-matching-molecules (dataset/).… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_train.moses_test_scaffolds_ours
nico8771/moses_test_scaffolds_ours — cleaned MOSES (test_scaffolds split)
Each row is a molecule as canonical SMILES plus RDKit-recomputed targets. MOSES is
neutral by construction, so molecules are featurized over 7 atom types with no
formal charges, and aromatic bonds are kept as their own class (no kekulization)
so the model learns aromaticity directly.
Source: official MOSES test_scaffolds.csv.gz (molecularsets/moses). Code:… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_test_scaffolds_ours.moses_test_defog
nico8771/moses_test_defog — MOSES (test) via DeFoG pipeline
Produced by DeFoG's exact moses_dataset.py process + filter_dataset=True
(build_molecule_with_partial_charges -> mol2smiles, single-fragment check), symbol-only
over 8 atom types with aromatic bonds, no explicit H. Molecules whose graph fails to
rebuild to a valid single-fragment SMILES (~10.5%, all pyrrole/imidazole [nH] kekulize
failures) are dropped — so this is the DeFoG filtered survivor set. Targets
(logP, qed… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_test_defog.moses_test_scaffolds_defog
nico8771/moses_test_scaffolds_defog — MOSES (test_scaffolds) via DeFoG pipeline
Produced by DeFoG's exact moses_dataset.py process + filter_dataset=True
(build_molecule_with_partial_charges -> mol2smiles, single-fragment check), symbol-only
over 8 atom types with aromatic bonds, no explicit H. Molecules whose graph fails to
rebuild to a valid single-fragment SMILES (~10.5%, all pyrrole/imidazole [nH] kekulize
failures) are dropped — so this is the DeFoG filtered survivor set.… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_test_scaffolds_defog.translation_dataset
Loading and Splitting Dataset To Various Languages
In this example, I will show you how to load the dataset and split by language for your downstream task.
>>> from datasets import load_dataset
>>> # load dataset
>>> dataset = load_dataset("mosesdaudu/translation_dataset")
>>> dataset
DatasetDict({
train: Dataset({
features: ['english_text', 'language', 'translated_text', 'split'],
num_rows: 198084
})
test: Dataset({
features: ['english_text'… See the full description on the dataset page: https://huggingface.co/datasets/mosesdaudu/translation_dataset.llava_instruction_80k
Note
this dataset is translated from llava_instruct_80k.json and has not been manually verified.
refer to images and dowload img datasets
JRC-Acquis_subcorpus_EL-FR_Hunalign_aligned-Moses
[!NOTE]
Dataset origin: https://inventory.clarin.gr/corpus/586
Description
The JRC-Acquis subcorpus EL-FR (Hunalign aligned-Moses) is a parallel subcorpus for French and Greek, subset of the JRC-Acquis Multilingual Parallel Corpus.
Citation
Joint Research Centre - European Commission (2015). JRC-Acquis subcorpus EL-FR (Hunalign aligned-Moses). [Dataset (Text corpus)]. CLARIN:EL. http://hdl.handle.net/11500/CLARIN-EL-0000-0000-69C4-D
moses_test_ours
nico8771/moses_test_ours — cleaned MOSES (test split)
Each row is a molecule as canonical SMILES plus RDKit-recomputed targets. MOSES is
neutral by construction, so molecules are featurized over 7 atom types with no
formal charges, and aromatic bonds are kept as their own class (no kekulization)
so the model learns aromaticity directly.
Source: official MOSES test.csv.gz (molecularsets/moses). Code:
https://github.com/Nico-Conti/flow-matching-molecules (dataset/).… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_test_ours.tone-on-a-budget-assets
tone-on-a-budget - runtime assets
Support files for the reference-free tone_i2 metric and the Tone on a Budget evaluation.
Code: https://github.com/mosesdaudu001/tone-on-a-budget
file
role
load in
f0_abs_calibration.v1.json
I2 speaker-normalised H/M/L decision boundaries
tone_f0_abs.py
probe/tone_probe_*
trained AfriHuBERT tone-probe weights (I1)
tone_probe.py
holdouts.v1.json
200 held-out Yoruba evaluation texts
notebooks
Note: holdouts.v1.json texts derive… See the full description on the dataset page: https://huggingface.co/datasets/mosesdaudu/tone-on-a-budget-assets.moses_test_scaffolds
nico8771/moses_test_scaffolds — cleaned MOSES (test_scaffolds split)
Each row is a molecule as canonical SMILES plus RDKit-recomputed targets. MOSES is
neutral by construction, so molecules are featurized over 7 atom types with no
formal charges, and aromatic bonds are kept as their own class (no kekulization)
so the model learns aromaticity directly.
Source: official MOSES test_scaffolds.csv.gz (molecularsets/moses). Code:
https://github.com/Nico-Conti/flow-matching-molecules… See the full description on the dataset page: https://huggingface.co/datasets/nico8771/moses_test_scaffolds.khmer-asr-datasetLegalLLMHK
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/MosesTan281/LegalLLMHK.
