CoolFace
23 results

auth

AuthenticIlm /Shamela4_Full_DB Shamela 4 — Full Islamic Library Corpus A complete extraction of al-Maktaba al-Shamela (الشاملة) v4, containing 8,589 books across 40 categories of classical Islamic sciences. Extracted from the original Lucene + Sqlite Shamela DB on 2026-04-26 with ~7.6 million pages and ~19 GB of Arabic text. Dataset Structure stage0_raw/ ├── _meta/ # Cross-cutting metadata (Parquet + JSONL) │ ├── extraction_manifest.json # Global extraction record │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/AuthenticIlm/Shamela4_Full_DB.text-generation10M<n<100M28 likes14k downloads4mo agoHugging Facekalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes4.6k downloads1y agoHugging Facelasrprobegen /authority-activationstext100K<n<1M0 likes2.7k downloads10mo agoHugging Facemainakmanna /single-author-arxiv Single-author arXiv Computer Science Metadata for arXiv records classified in Computer Science that list exactly one author. default retains the original daily-file import. fast stores historical data in monthly files and adds new submissions as daily update files; it is the configuration used by the public archive because it makes filtering much faster. This dataset contains metadata only. arXiv is the source of truth; use each record's arxiv_url and pdf_url to read the paper. text100K<n<1M0 likes2.3k downloads1mo agoHugging Facebarilan /blog_authorship_corpusThe Blog Authorship Corpus consists of the collected posts of 19,320 bloggers gathered from blogger.com in August 2004. The corpus incorporates a total of 681,288 posts and over 140 million words - or approximately 35 posts and 7250 words per person. Each blog is presented as a separate file, the name of which indicates a blogger id# and the blogger’s self-provided gender, age, industry and astrological sign. (All are labeled for gender and age but for many, industry and/or sign is marked as unknown.) All bloggers included in the corpus fall into one of three age groups: - 8240 "10s" blogs (ages 13-17), - 8086 "20s" blogs (ages 23-27), - 2994 "30s" blogs (ages 33-47). For each age group there are an equal number of male and female bloggers. Each blog in the corpus includes at least 200 occurrences of common English words. All formatting has been stripped with two exceptions. Individual posts within a single blogger are separated by the date of the following post and links within a post are denoted by the label urllink. The corpus may be freely used for non-commercial research purposes.text-classification10K<n<100K18 likes1.1k downloads3y agoHugging Facehkadxqq /spooky-author-identificationtext10K<n<100K0 likes981 downloads4y agoHugging Face