forums
Datasets
All datasets matching “forums”Roleplay-Forums_2023-04Important noticeUpon reanalysis, it appears that spaces between adjacent HTML tags in these scrapes got mangled, a problem that was also present in the files of the Rentry where they first got uploaded; this occurred during an intermediate step in the conversion process from the original source files, most of which unfortunately do not exist anymore. So, the usefulness of the data will be diminished. Some scrapes without these issues have been provided on a different dataset page.… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Roleplay-Forums_2023-04.roleplaying-forums-raw
Roleplaying forum scrapes (raw)
Here are mostly original/raw files for some of the roleplaying forums I scraped in the past (and some newly scraped ones), repacked as HTML strings + some metadata on a one-row-per-thread basis instead of a one-row-per-message basis, which should make them more convenient to handle.
Unlike the previously uploaded archive, they shouldn't have issues with spaces between adjacent HTML tags, as that occurred by mistake in an intermediate processing step… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/roleplaying-forums-raw.forums_pol_json_zstforumsoh_tokenized
ForumSohbetleri Tokenized (All Subsets)
This dataset is a pre-tokenized, shuffled, and interleaved version of all subsets from turkish-nlp-suite/ForumSohbetleri.
Processing Details
Tokenizer: Ba2han/qwen-test-3
Filtering: Min tokens = 50, Max tokens = 2550
Format: uint32 input_ids only.
Subsets Included: donanimarsivi, donanimhaber, forumum, iyinet, kadinlarklubu, memurlar, tahribat, technopatsosyal, turkiyeforum, wardom, wmaraci
Shuffling: Stream interleaved with a… See the full description on the dataset page: https://huggingface.co/datasets/Ba2han/forumsoh_tokenized.lemonilia-Roleplay-Forums
Roleplay Forums 2023-04
As part of a probably misguided effort for gathering useful data for a roleplay finetune, around April 2023 I scraped several of the largest or most popular English-language roleplaying forums publicly known, for a total of about 47 GB of data (uncompressed, including HTML tags and metadata), and posted the files on a public Rentry page. I handpicked and cleaned a tiny fraction of that data for LimaRP, among other things. Other people posted partially cleaned… See the full description on the dataset page: https://huggingface.co/datasets/NewEden-Forge/lemonilia-Roleplay-Forums.ForumSohbetleri
Dataset Card for ForumSohbetleri
ForumSohbetleri a web forum tetx corpus for Turkish, indeed first large-scale Turkish forum text corpus.
This corpus is a part of large scale Turkish corpus Bella Turca. For more details about Bella Turca, please refer to the publication.
This collection is made up of several subsets, each subset is gathered from the corresponding forum website. Forum websites contains diverse topics, ladies only, tech, economics, life, relations and much more...… See the full description on the dataset page: https://huggingface.co/datasets/turkish-nlp-suite/ForumSohbetleri.
