CoolFace
20 results

farsi

Thomcles /Persian-Farsi-Speechgated Persian (Farsi) TTS Dataset 🗂️ Dataset Description This dataset is a Persian (Farsi) text-to-speech (TTS) corpus built by concatenating and denoising multiple existing Farsi datasets.It is intended for training and evaluation of speech synthesis (TTS) models in Persian. Since the basic datasets were contaminated with unintelligible audio, I used dnsmos to keep only clean audio (mos_ovr >= 3.0, same value as for the Emilia dataset). The dataset contains two main… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/Persian-Farsi-Speech.audiotext-to-speech100K<n<1M23 likes1.9k downloads29d agoHugging FaceParsiAI /FarsInstruct News [2025.01.20] 🏆 Our paper was nominated as the best paper at LowResLM @ COLING 2025! [2024.12.07] ✨ Our paper has been accepted for oral presentation at LowResLM @ COLING 2025! Dataset Summary Instruction-tuned large language models have demonstrated remarkable capabilities in following human instructions across various domains. However, their proficiency remains notably deficient in many low-resource languages. To address this challenge, we begin by… See the full description on the dataset page: https://huggingface.co/datasets/ParsiAI/FarsInstruct.texttext-classification10M<n<100M23 likes1.7k downloads2y agoHugging FacekiarashQ /farsi-asr-unified-cleaned 🎧 Farsi ASR Unified Dataset (Parquet Sharded Edition) Overview The Farsi ASR Unified Dataset is a large-scale, high-quality, and fully standardized collection of Persian (Farsi) speech-to-text data — designed specifically for modern machine learning and ASR (Automatic Speech Recognition) workflows. This dataset consolidates audio–text pairs from multiple open sources, applies a rigorous cleaning and normalization pipeline, and stores everything efficiently in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/kiarashQ/farsi-asr-unified-cleaned.audio1M<n<10M6 likes1.5k downloads11mo agoHugging FacePeacockery /farsi-asr-iran-international-raw Iran International raw Farsi audio archive This repository preserves 11,993 individually addressable FLAC source files for incremental ASR relabeling and reproducible restoration. Repository file layout The first 9,990 FLAC files are stored at the repository root. The remaining 2,003 FLAC files are stored individually under overflow/ to respect Hugging Face's 10,000-entry-per-directory limit. REMOTE_PATHS.jsonl records every source filename, remote path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-iran-international-raw.audio10K<n<100K1 likes1.3k downloads2mo agoHugging Facelinks-ads /farsite-layers PanEU FARSITE STAC Catalog Dataset Description This dataset provides harmonised, continent-wide raster layers at approximately 74m spatial resolution covering Europe, developed under the FIRE‑RES programme by the CIRGEO Centre at the University of Padova. It includes information on surface fuel models, canopy fuel attributes (such as canopy height, canopy cover, bulk density), and topographic features. These layers are co-registered and designed to support… See the full description on the dataset page: https://huggingface.co/datasets/links-ads/farsite-layers.imagen<1K0 likes650 downloads11mo agoHugging Facesrezas /farsi_voice_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/srezas/farsi_voice_dataset.audioautomatic-speech-recognition100K<n<1M5 likes511 downloads2y agoHugging Face