karaka
KaraKaraWitch-L3.1-70b-Swallow-Saigetsu-GGUFKaraKaraWitch-Llama-ProgressPushDoll-3.3-70Bees-GGUFKaraKaraWitch-L3.1-70b-Inori-GGUFKaraKaraWitch-Llama-3.X-Workout-70B-GGUFKaraKaraWitch-LLENN-v0.75-Qwen2.5-72b-GGUFKaraKaraWitch-Llama-MiraiFanfare-2-3.3-70B-GGUFKaraKaraWitch_-_L3.1-70b-Inori-ggufKaraKaraWitch-Llama-MiraiFanfare-3.3-70B-GGUF
Datasets
All datasets matching “karaka”AnimuSubtitle-JP
KaraKaraWitch/AnimuSubtitle-JP
This dataset is an extract of Subtitles from anime. The original files are sourced from nyaa.si.
Folders
data (Deprecated)
data_ass (Extracted .ass subtitles from <SOURCE A>)
data_TS (Extracted arib subtitles from <SOURCE B>)
Dataset Format
Dataset is in Advanced SubStation Alpha (colloquially known as as ASS) SSA/ASS Specs.
If you're looking for programmatic access, you may parse the file with this python library.
import ass… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/AnimuSubtitle-JP.MyselfAndEveryone
KaraKaraWitch/MyselfAndEveryone
Might as well have a dataset card.
This is a dataset dumping ground. Stuff that I'm working on or in the works that I want to upload and share.
I think most of the stuff is encrypted but there's some that are not. If it's encrypted, it's for someone or myself and not exactly for public use.
karakalpak-speech-corpus
📚 Karakalpak Speech Corpus (107 Hours)
The Karakalpak Speech Corpus is the first comprehensive, open-access, community-crowdsourced speech recognition dataset for the Karakalpak language (kaa), a low-resource Turkic language spoken primarily in the Republic of Karakalpakstan (Uzbekistan).
Founded and led by Atabek Kadirbergenov alongside a student research team from the Muhammad al-Khwarizmi Specialized School in Nukus, this dataset was created to preserve cultural heritage… See the full description on the dataset page: https://huggingface.co/datasets/atikuwu/karakalpak-speech-corpus.karakaijouzunotakagisan3
Bangumi Image Base of Karakai Jouzu No Takagi-san 3
This is the image base of bangumi Karakai Jouzu no Takagi-san 3, we detected 34 characters, 5592 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/karakaijouzunotakagisan3.TvTroper-2025
TvTroper-2025
A cleaned & refreshed dump of ~708 k pages from tvtropes.org
Dataset Summary
TvTroper-2025 is an updated snapshot of TvTropes.org (≈ 708 000 wiki pages, namespaces and date-grouped pages excluded).
Every page is released in two flavours:
Raw HTML – 22 GB single file
Markdown-cleaned – split into 1 GB JSONL shards (no unpacking required)
No additional content filtering has been applied; short sub-index pages are left in so you can decide what to drop.… See the full description on the dataset page: https://huggingface.co/datasets/KaraKaraWitch/TvTroper-2025.karakaijouzunotakagisan
Bangumi Image Base of Karakai Jouzu No Takagi-san
This is the image base of bangumi Karakai Jouzu no Takagi-san, we detected 21 characters, 6297 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1%… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/karakaijouzunotakagisan.
