CoolFace
Datasetpublic

trentmkelly/asianfanfics

Asianfanfics Archive This dataset contains metadata and public chapter text from Asianfanfics, packaged as zstd-compressed Parquet files. Repository: trentmkelly/asianfanfics Contents The dataset is split into two Hugging Face configs: Table Rows Files Description stories 1,727,491 4 Story-level metadata and text-fetch status. chapters 4,114,576 42 Chapter-level metadata, fetch status, and fetched public text when available. Total compressed… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/asianfanfics.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
1likes151downloads
Dataset Card

Asianfanfics Archive

This dataset contains metadata and public chapter text from Asianfanfics, packaged as zstd-compressed Parquet files.

Repository: trentmkelly/asianfanfics

Contents

The dataset is split into two Hugging Face configs:

TableRowsFilesDescription
stories1,727,4914Story-level metadata and text-fetch status.
chapters4,114,57642Chapter-level metadata, fetch status, and fetched public text when available.

Total compressed Parquet size: about 4.2 GiB.

Text Coverage

Chapter fetch status counts:

StatusRowsMeaning
fetched2,045,265Public chapter text is present in chapters.text.
skipped_gated1,236,406Chapter/story appeared login-, member-, or subscription-gated.
missing_body820,399No static public chapter body was available.
empty_body11,776Chapter body was present but empty.
not_found730Chapter page was unavailable.

Stored chapter text totals approximately 14,918,091,815 characters.

Story text status counts:

StatusStories
fetched302,857
no_public_text235,124
skipped_gated211,049

Unavailable story IDs are included in stories for index completeness, but do not have story text.

Columns

stories

Includes story ID, URL, title, author, status/rating, publication/update metadata, engagement counters, tags/characters JSON, availability status, and text-fetch status fields.

chapters

Includes story ID, chapter index, URL, title, word-count metadata, fetch status, fetch error, and text for chapters where public text was fetched.

License

This dataset is released under the Creative Commons Attribution-ShareAlike 4.0 International license (CC-BY-SA-4.0).