datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.Malaysian-TTS-v2
Malaysian TTS v2
Generate Malay and localize English for TTS dataset, currently only support 2 speakers, husein and idayu, where total audio is 4642.77 hours.
How to prepare the dataset
huggingface-cli download \
mesolitica/Malaysian-TTS-v2 \
--include "all-*.zip" \
--repo-type "dataset" \
--local-dir './'
huggingface-cli download \
mesolitica/STT-Normalizer \
--include "*husein*.zip" \
--exclude "*force*" \
--repo-type "dataset" \
--local-dir './'… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS-v2.fineweb-filter-malaysian-context
HuggingFaceFW/fineweb filter Malaysian context
What is it?
We filter the original 🍷 FineWeb dataset that consists more than 15T tokens on simple Malaysian keywords.
Total tokens for the filtered dataset is 174102784199 tokens, 174B tokens.
How we do it?
We filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords on r5.16xlarge EC2 instance for 7 days.
We calculate total tokens using tiktoken.encoding_for_model("gpt2") on c7a.24xlarge EC2… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/fineweb-filter-malaysian-context.Malaysian-Emilia-v2
Malaysian Emilia v2
This version 2 should fixed https://github.com/open-mmlab/Amphion/issues/436, an Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Malaysian and Singaporean Speech Generation. Replicating Emilia on,
Dataset
Clone and Extract
We upload as split zip files so you can clone and extract distributedly,
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2.MesoMathematics
MesoMathematics
Frozen data artifacts for the paper “Mathematical Knowledge at the
Mesoscale: Organization after the Formal Mathematics Revolution” by Andrea
E. V. Ferrari, Benjy Firester, Xinze Li, Simone Severini, and Patrick Shafto.
The corresponding source code, exact commands, and manuscript live in the
MathNetwork/MesoMathematics
repository.
This release fixes Mathlib at v4.33.0, commit
db584cd6d46c92f209a44c0f1c829460d327499d, with Lean v4.33.0. The full
commit, not the… See the full description on the dataset page: https://huggingface.co/datasets/CloKTech/MesoMathematics.tcga-meso-tabular-open
TCGA-MESO — Tabular (Open Access)
Open-access TCGA-MESO data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV.
GDC data release: Data Release 46.0 - August 10, 2026
Built: 2026-09-12 04:10:59 UTC
Scope: one TCGA project — see [the family][repo] for the others
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-meso-tabular-open.pseudolabel-malaya-speech-stt-train-whisper-large-v3TTS-Combine-annotated
Replicating HuggingFace Dataspeech using Malay dataset
This is combination of mesolitica/tts-azure-annotated and mesolitica/tts-gtts-annotated
Speakers
Yasmin, ID 0, female
Osman, ID 1, male
Bunga, ID 2, female
Ariff, ID 3, male
Ayu, ID 4, female
Kamarul, ID 5, male
Danial, ID 6, male
Elina, ID 7, female
With total ~713 hours.
Source code
Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/text-to-speech/dataspeech
africa-synth-mental-health-asbestos-mesothelioma-all
Asbestos Exposure & Mesothelioma (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mental-health-asbestos-mesothelioma-all.mesopotamia
CITATION:
If you use the dataset kindly cite our paper DomAINS - DOMain Adapted INStructions.
example_dataset
example_dataset
This dataset was generated using phosphobot.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot.
To get started in robotics, get your own phospho starter pack..
Bausstest_lagi_20260616_225943This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/abbas-mesolitica/Bausstest_lagi_20260616_225943.Bausstest_lelab_20260616_225213This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/abbas-mesolitica/Bausstest_lelab_20260616_225213.
