CoolFace
20 results

karaka

KaraKaraWitch /AnimuSubtitle-JP KaraKaraWitch/AnimuSubtitle-JP This dataset is an extract of Subtitles from anime. The original files are sourced from nyaa.si. Folders data (Deprecated) data_ass (Extracted .ass subtitles from <SOURCE A>) data_TS (Extracted arib subtitles from <SOURCE B>) Dataset Format Dataset is in Advanced SubStation Alpha (colloquially known as as ASS) SSA/ASS Specs. If you're looking for programmatic access, you may parse the file with this python library. import ass… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/AnimuSubtitle-JP.text-classification4 likes922 downloads2y agoHugging FaceKaraKaraWitch /MyselfAndEveryone KaraKaraWitch/MyselfAndEveryone Might as well have a dataset card. This is a dataset dumping ground. Stuff that I'm working on or in the works that I want to upload and share. I think most of the stuff is encrypted but there's some that are not. If it's encrypted, it's for someone or myself and not exactly for public use. 2 likes837 downloads2y agoHugging Faceatikuwu /karakalpak-speech-corpus 📚 Karakalpak Speech Corpus (107 Hours) The Karakalpak Speech Corpus is the first comprehensive, open-access, community-crowdsourced speech recognition dataset for the Karakalpak language (kaa), a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan (Uzbekistan). Founded and led by Atabek Kadirbergenov alongside a student research team from the Muhammad al-Khwarizmi Specialized School in Nukus, this dataset was created to preserve cultural heritage… See the full description on the dataset page: https://huggingface.co/datasets/atikuwu/karakalpak-speech-corpus.textautomatic-speech-recognition10K<n<100K5 likes595 downloads3d agoHugging FaceBangumiBase /karakaijouzunotakagisan3 Bangumi Image Base of Karakai Jouzu No Takagi-san 3 This is the image base of bangumi Karakai Jouzu no Takagi-san 3, we detected 34 characters, 5592 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/karakaijouzunotakagisan3.image1K<n<10K0 likes579 downloads2y agoHugging FaceKaraKaraWitch /TvTroper-2025 TvTroper-2025 A cleaned & refreshed dump of ~708 k pages from tvtropes.org Dataset Summary TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded). Every page is released in two flavours: Raw HTML – 22 GB single file Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required) No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.texttext-classification1M<n<10M4 likes419 downloads11mo agoHugging FaceBangumiBase /karakaijouzunotakagisan Bangumi Image Base of Karakai Jouzu No Takagi-san This is the image base of bangumi Karakai Jouzu no Takagi-san, we detected 21 characters, 6297 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/karakaijouzunotakagisan.image1K<n<10K0 likes388 downloads3y agoHugging Face