CoolFace
20 results

punctuation

ai-forever /spellcheck_punctuation_benchmarkRussian Spellcheck Benchmark is a new benchmark for spelling correction in Russian language. It includes four datasets, each of which consists of pairs of sentences in Russian language. Each pair embodies sentence, which may contain spelling errors, and its corresponding correction. Datasets were gathered from various sources and domains including social networks, internet blogs, github commits, medical anamnesis, literature, news, reviews and more.text-generation10K<n<100K5 likes412 downloads2y agoHugging Facegovnejri /golos_mfa_punctuation Golos MFA Punctuation Расширенная версия датасета Golos — русскоязычного корпуса речи с краудсорс и студийными записями. Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Русский (ru) Записей 970,597 Аудио ~1,044 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Что добавлено по сравнению с оригинальным Golos… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/golos_mfa_punctuation.audio100K<n<1M5 likes402 downloads5mo agoHugging Facegovnejri /kazakh_speech_mfa_punctuation Kazakh Speech MFA Punctuation Расширенная версия датасета ISSAI KSC2 — крупнейшего открытого корпуса казахской речи от института ISSAI (Nazarbayev University). Датасет дополнен пунктуацией и word-level временными метками (MFA alignment). Опубликовано и поддерживается Jeti Labs. Описание Параметр Значение Язык Казахский (kk) Записей 595,690 Аудио ~1,110 часов Частота дискретизации 16,000 Hz Формат WAV, mono, 16-bit Размер 52.9 GB… See the full description on the dataset page: https://huggingface.co/datasets/govnejri/kazakh_speech_mfa_punctuation.audio100K<n<1M6 likes178 downloads1mo agoHugging Faceadalbertojunior /punctuation-ptbrtext100K<n<1M0 likes154 downloads5y agoHugging FaceGoktugD /turkish-punctuation-restoration-500k Turkish Punctuation Restoration 500K v2 Noktalama ve büyük harfleri kaldırılmış girişler ile hedef cümle çiftleri. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, unpunctuated_text, punctuated_text Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-punctuation-restoration-500k.texttext-generation100K<n<1M0 likes146 downloads1mo agoHugging Faceadalbertojunior /punctuation-ptbr-lighttext10K<n<100K0 likes130 downloads5y agoHugging Face