CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SilencioNetwork /cebuano-speech Cebuano (Bisaya) Spontaneous Speech — Silencio Philippines Pack Spontaneous long-form Cebuano with human transcription and word-level forced alignment. Fifteen speakers, mean clip length over two minutes, 27,000+ timestamped tokens. Part of the Silencio Philippines Pack. Hours 3.48 Clips 90 Speakers 15 Countries 2 Speaker origin regions 4 L1 speakers of the recorded language 11 of 15 (65 clips) Audio 48 kHz stereo WAV Mean clip length 139.2 s… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/cebuano-speech.audioautomatic-speech-recognitionn<1K0 likes288 downloads4d agoHugging Face02Jession01 /English-Cebuano-Translation texttext-generation100K<n<1M0 likes49 downloads3y agoHugging Face03filbench /cebuano-readabilitySource: https://github.com/imperialite/cebuano-readability We asked permission from one of the authors to include this dataset to our catalog effort. We copy a portion of the README in this dataset card. Baseline Readability Assessment Model for Cebuano This repository contains the code and datasets from Bloom, Let's Read Asia, and Department of Education (DepEd) websites used for developing the first ML-based baseline for readability assessment in the Cebuano language described… See the full description on the dataset page: https://huggingface.co/datasets/filbench/cebuano-readability.textn<1K0 likes41 downloads2y agoHugging Face04Speech-data /Cebuano-Speech-Dataset 🎧 Cebuano Speech Dataset The Cebuano Speech Dataset is a high-quality speech audio dataset designed to deliver structured and diverse audio data for AI-powered voice applications. It includes 108 hours of audio data distributed across 807 files, provided in MP3 and WAV formats, with a total size of 135 MB. This well-organized audio dataset ensures balanced voice data, with 49% female and 51% male speakers, and a broad age range from 18 to 50+ years. The dataset language is Cebuano… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Cebuano-Speech-Dataset.audioautomatic-speech-recognitionn<1K1 likes28 downloads6mo agoHugging Face05jojo-ai-mst /Roleplay-Cebuano RolePlay-Cebuano Roleplay-Cebuano Dataset is a dataset for roleplaying in the Amharic language for the Large Language Model. The base dataset is the GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API. For more information and other language datasets for roleplay, see this github repo. For… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Cebuano.texttext-generation1K<n<10K0 likes26 downloads2y agoHugging Face06jfernandez /cebuano-filipino-sentencestext100K<n<1M4 likes21 downloads4y agoHugging Face07ljvmiranda921 /gsd-translate-Cebuanotext1K<n<10K0 likes20 downloads3mo agoHugging Face08ljvmiranda921 /gsd-smith-Cebuanotext1K<n<10K0 likes18 downloads4mo agoHugging Face09saillab /alpaca_cebuano_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_cebuano_taco.text10K<n<100K1 likes17 downloads2y agoHugging Face10mikhail-panzo /cebuano-processed1K<n<10K0 likes15 downloads2y agoHugging Face11roncc13 /CebuanoAnnotatedFakeandLegitNewstext1K<n<10K0 likes13 downloads3mo agoHugging Face12ljvmiranda921 /gsd-teacher-Cebuanotext1K<n<10K0 likes12 downloads3mo agoHugging Face13youdiniplays /tagalog-cebuano_translationtext100K<n<1M1 likes11 downloads3y agoHugging Face14ljvmiranda921 /gsd-judgelm-traj-Cebuanotextn<1K0 likes11 downloads3mo agoHugging Face15saillab /alpaca-cebuano-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-cebuano-cleaned.text10K<n<100K0 likes9 downloads2y agoHugging Face16ECENYAS /Filipino_Cebuanotextn<1K0 likes8 downloads10mo agoHugging Face17jelovrgr /Cebuano_translation_datatext10K<n<100K0 likes7 downloads2y agoHugging Face18ECENYAS /Tagalog_Cebuanotextn<1K0 likes7 downloads1y agoHugging Face19e-SALIN /Tagalog_to_Cebuano_mBARTtextn<1K0 likes6 downloads2y agoHugging Face20nielle003 /Cebuano_Iloko_Tagalog_Question-Answerquestion-answering0 likes6 downloads6mo agoHugging Face21e-SALIN /Tagalog_to_Cebuano_marianMTtextn<1K0 likes5 downloads2y agoHugging Face22minnalproc /cebuano_affix_listtextn<1K0 likes4 downloads6mo agoHugging Face23ECENYAS /Filipino_to_Cebuano0 likes3 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.