CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mesolitica /Malaysian-Emilia-annotated Malaysian Emilia Annotated Annotate Malaysian-Emilia using Data-Speech pipeline. Malaysian Youtube Originally from malaysia-ai/crawl-youtube Total 3168.8 hours. Gender prediction, filtered-24k_processed_24k_gender.zip Language prediction, filtered-24k_processed_language.zip Force alignment. Post cleaned to 24k and 44k sampling rates, 24k, filtered-24k_processed_24k.zip 44k, filtered-24k_processed_44k.zip Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.tabulartext-to-speech1M<n<10M2 likes2.1k downloads1y agoHugging Face02mesolitica /Malaysian-TTS-v2 Malaysian TTS v2 Generate Malay and localize English for TTS dataset, currently only support 2 speakers, husein and idayu, where total audio is 4642.77 hours. How to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-TTS-v2 \ --include "all-*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/STT-Normalizer \ --include "*husein*.zip" \ --exclude "*force*" \ --repo-type "dataset" \ --local-dir './'… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-TTS-v2.tabular1M<n<10M2 likes1.2k downloads1y agoHugging Face03mesolitica /fineweb-filter-malaysian-context HuggingFaceFW/fineweb filter Malaysian context What is it? We filter the original 🍷 FineWeb dataset that consists more than 15T tokens on simple Malaysian keywords. Total tokens for the filtered dataset is 174102784199 tokens, 174B tokens. How we do it? We filter rows using {'malay', 'malaysia', 'melayu', 'bursa', 'ringgit'} keywords on r5.16xlarge EC2 instance for 7 days. We calculate total tokens using tiktoken.encoding_for_model("gpt2") on c7a.24xlarge EC2… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/fineweb-filter-malaysian-context.tabular10M<n<100M1 likes804 downloads2y agoHugging Face04Scicom-intl /Malaysian-Emilia Malaysian Emilia Gather Malaysian Emilia from, https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2 https://huggingface.co/datasets/Scicom-intl/Malaysian-Chinese-Emilia https://huggingface.co/datasets/mesolitica/Malaysian-Emilia#malaysian-dialect And do, Trim silent. Permutation for Voice Conversion include post-filtering during permutation. Convert to Neucodec speech tokens. tabular10M<n<100M1 likes654 downloads7mo agoHugging Face05Scicom-intl /Malaysian-Chinese-Emilia Malaysian-Chinese-Emilia Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Chinese audio. Total rows: 605169 Total hours: 1857.611445057867 hours Permutation for Voice Conversion Also we already calculated speaker permutation to prepare for voice conversion. tabular10M<n<100M1 likes614 downloads8mo agoHugging Face06mesolitica /Malaysian-Emilia-v2 Malaysian Emilia v2 This version 2 should fixed https://github.com/open-mmlab/Amphion/issues/436, an Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Malaysian and Singaporean Speech Generation. Replicating Emilia on, Dataset Clone and Extract We upload as split zip files so you can clone and extract distributedly, huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2.tabular1M<n<10M2 likes310 downloads1y agoHugging Face07Scicom-intl /Malaysian-RAG-QA Malaysian RAG Dataset This dataset contains question-answering pairs with associated context documents and evaluation metrics. Each entry includes a source document, a question and an answer generated by ChatGPT-5.1, and quality scores for context precision and faithfulness evaluated using RAGAS. Data Sources Documents sourced from malaysia-ai/pretrain-text-dataset: maktabahalbakri.com muftiwp.gov.my.dedup asklegal dewanbahasa-jdbp gov.my Dataset Splits… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Malaysian-RAG-QA.tabular1K<n<10K0 likes32 downloads9mo agoHugging Face08skilledu /Malaysian-TTS-v2 Malaysian TTS v2 Generate Malay and localize English for TTS dataset, currently only support 2 speakers, husein and idayu, where total audio is 4642.77 hours. How to prepare the dataset huggingface-cli download \ mesolitica/Malaysian-TTS-v2 \ --include "all-*.zip" \ --repo-type "dataset" \ --local-dir './' huggingface-cli download \ mesolitica/STT-Normalizer \ --include "*husein*.zip" \ --exclude "*force*" \ --repo-type "dataset" \ --local-dir './'… See the full description on the dataset page: https://huggingface.co/datasets/skilledu/Malaysian-TTS-v2.tabular1M<n<10M0 likes31 downloads4mo agoHugging Face09Scicom-intl /Malaysian-Emilia-Sidon Malaysian-Emilia-Sidon Apply sarulab-speech/sidon-v0.1 on, https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2 https://huggingface.co/datasets/Scicom-intl/Malaysian-Chinese-Emilia https://huggingface.co/datasets/Scicom-intl/Malaysian-Tamil-Emilia tabular1M<n<10M0 likes18 downloads8mo agoHugging Face10Scicom-intl /Malaysian-Tamil-Emilia Malaysian-Tamil-Emilia Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Tamil audio. tabular10M<n<100M0 likes13 downloads7mo agoHugging Face11ilyass31 /MH370_Malaysian_Airlines_Satellite_Data MH370 Inmarsat Satellite Data Logs Overview This dataset contains the data communication logs from the Inmarsat satellite system related to Malaysia Airlines Flight MH370 (9M-MRO). The data was provided to the authorities to assist in the ongoing investigation and search efforts for the missing aircraft. Dataset Contents The dataset consists of structured logs recorded at the Ground Earth Station (GES), capturing communications between the Inmarsat satellite… See the full description on the dataset page: https://huggingface.co/datasets/ilyass31/MH370_Malaysian_Airlines_Satellite_Data.tabulartext-classificationn<1K2 likes7 downloads2y agoHugging Face12Scicom-intl /Malaysian-Call-Center-Language-Switchingtabular1K<n<10K0 likes5 downloads4mo agoHugging Face13userdata /filtered-pseudolabel-malaysian-youtube-whisper-large-v3tabular100K<n<1M0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.