CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes2.1k downloads1y agoHugging Face02mesolitica /Malaysian-TTS-v2 Malaysian TTS v2 Generate Malay and localize English for TTS dataset, currently only support 2 speakers, husein and idayu, where total audio is 4642.77 hours. How to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-TTS-v2 \ --include "all-*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/STT-Normalizer \ --include "*husein*.zip" \ --exclude "*force*" \ --repo-type "dataset" \ --local-dir './'… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS-v2.tabular1M<n<10M2 likes1.2k downloads1y agoHugging Face03mesolitica /fineweb-filter-malaysian-context HuggingFaceFW/fineweb filter Malaysian context What is it? We filter the original 🍷 FineWeb dataset that consists more than 15T tokens on simple Malaysian keywords. Total tokens for the filtered dataset is 174102784199 tokens, 174B tokens. How we do it? We filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords on r5.16xlarge EC2 instance for 7 days. We calculate total tokens using tiktoken.encoding_for_model("gpt2") on c7a.24xlarge EC2… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/fineweb-filter-malaysian-context.tabular10M<n<100M1 likes804 downloads2y agoHugging Face04mesolitica /Malaysian-Emilia-v2 Malaysian Emilia v2 This version 2 should fixed https://github.com/open-mmlab/Amphion/issues/436, an Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Malaysian and Singaporean Speech Generation. Replicating Emilia on, Dataset Clone and Extract We upload as split zip files so you can clone and extract distributedly, huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2.tabular1M<n<10M2 likes310 downloads1y agoHugging Face05CloKTech /MesoMathematics MesoMathematics Frozen data artifacts for the paper “Mathematical Knowledge at the Mesoscale: Organization after the Formal Mathematics Revolution” by Andrea E. V. Ferrari, Benjy Firester, Xinze Li, Simone Severini, and Patrick Shafto. The corresponding source code, exact commands, and manuscript live in the MathNetwork/MesoMathematics repository. This release fixes Mathlib at v4.33.0, commit db584cd6d46c92f209a44c0f1c829460d327499d, with Lean v4.33.0. The full commit, not the… See the full description on the dataset page: https://huggingface.co/datasets/CloKTech/MesoMathematics.tabular100K<n<1M0 likes240 downloads1mo agoHugging Face06gabrielaltay /tcga-meso-tabular-open TCGA-MESO — Tabular (Open Access) Open-access TCGA-MESO data from the NCI Genomic Data Commons, reshaped into one table per GDC data_type. Clinical, biospecimen and every open molecular modality for this cohort, in one place, queryable without downloading a single .tar or parsing a single TSV. GDC data release: Data Release 46.0 - August 10, 2026 Built: 2026-09-12 04:10:59 UTC Scope: one TCGA project — see [the family][repo] for the others from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/gabrielaltay/tcga-meso-tabular-open.tabular10M<n<100M1 likes216 downloads13d agoHugging Face07mesolitica /pseudolabel-malaya-speech-stt-train-whisper-large-v3tabularautomatic-speech-recognition1M<n<10M1 likes63 downloads3y agoHugging Face08mesolitica /TTS-Combine-annotated Replicating HuggingFace Dataspeech using Malay dataset This is combination of mesolitica/tts-azure-annotated and mesolitica/tts-gtts-annotated Speakers Yasmin, ID 0, female Osman, ID 1, male Bunga, ID 2, female Ariff, ID 3, male Ayu, ID 4, female Kamarul, ID 5, male Danial, ID 6, male Elina, ID 7, female With total ~713 hours. Source code Notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/text-to-speech/dataspeech tabular100K<n<1M0 likes57 downloads1y agoHugging Face09electricsheepafrica /africa-synth-mental-health-asbestos-mesothelioma-all Asbestos Exposure & Mesothelioma (SSA) | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mental-health-asbestos-mesothelioma-all.imagetabular-classificationn<1K0 likes39 downloads1mo agoHugging Face10meithnav /mesopotamia CITATION: If you use the dataset kindly cite our paper DomAINS - DOMain Adapted INStructions. tabular100K<n<1M1 likes33 downloads1y agoHugging Face11abbas-mesolitica /example_dataset example_dataset This dataset was generated using phosphobot. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot. To get started in robotics, get your own phospho starter pack.. tabularroboticsn<1K0 likes31 downloads4mo agoHugging Face12abbas-mesolitica /Bausstest_lagi_20260616_225943This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/abbas-mesolitica/Bausstest_lagi_20260616_225943.tabularrobotics1K<n<10K0 likes30 downloads3mo agoHugging Face13abbas-mesolitica /Bausstest_lelab_20260616_225213This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/abbas-mesolitica/Bausstest_lelab_20260616_225213.tabularroboticsn<1K0 likes28 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.