CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mideind /icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3). texttext-generation1M<n<10M1 likes227 downloads4y agoHugging Face02amzar1303 /crawl-karangan-net-komsastext1K<n<10K0 likes141 downloads3y agoHugging Face03OpenTransformer /agillm-crawl-datatext100K<n<1M0 likes95 downloads3mo agoHugging Face04OpenTransformer /goddess-crawltext10M<n<100M0 likes92 downloads7mo agoHugging Face05ammarxix /crawl-lirik-lagu-dot-nettext1K<n<10K0 likes79 downloads3y agoHugging Face06acul3 /link_crawl_beritatext1M<n<10M0 likes77 downloads3y agoHugging Face07projecte-aina /catalan_government_crawling Dataset Card for Catalan Government Crawling Dataset Summary The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39,117,909 tokens, 1,565,433 sentences and 71,043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_government_crawling.textfill-mask10K<n<100K1 likes57 downloads2y agoHugging Face08macja /CRAwLeR-PL CRAwLeR-PL — Cross-Reference Aware Legal Retrieval (Polish) CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Polish instance, built on Polish legal Acts obtained through the ELI API. It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-PL.texttext-retrieval10K<n<100K0 likes57 downloads3mo agoHugging Face09hllj /vi_math_problem_crawl Dataset Card for Vietnamese Elementary Math Knowledge and Workbook Dataset Summary The data includes information about elementary school math knowledge in Vietnam, as well as exercises compiled from books. This is a crawlable dataset that can be trained for text generation tasks. Supported Tasks and Leaderboards Languages The majority of the data is in Vietnamese, but there is still some English from some bilingual workbooks. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_math_problem_crawl.texttext-generation10K<n<100K1 likes53 downloads3y agoHugging Face10macja /CRAwLeR-DK CRAwLeR-DK — Cross-Reference Aware Legal Retrieval (Danish) CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Danish instance, built on Danish legal documents from Retsinformation. It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-DK.texttext-retrieval10K<n<100K0 likes53 downloads3mo agoHugging Face11osamamumtaz01 /ai-crawler-user-agents AI Crawler User Agents Machine-readable list of all 28 known AI crawler and agent user-agent strings — GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider, Applebot-Extended, and more — with each bot's operator, purpose, observed robots.txt compliance, and documented crawl-delay support. Fields Field Description userAgent Token to match in robots.txt / server logs (e.g. GPTBot) operator Company running the crawler purpose… See the full description on the dataset page: https://huggingface.co/datasets/osamamumtaz01/ai-crawler-user-agents.textn<1K0 likes41 downloads9d agoHugging Face12aisyahhrazak /crawl-malaysian-websitetext10K<n<100K0 likes39 downloads3y agoHugging Face13wanadzhar913 /crawl-leaazleeya TLDR Website: leaazleeya Num. of webpages: 543 Num. of webpages scraped: 543 Num. articles successfully extracted: 534 Remaing webpages to be scraped: 0 Scraped on: 5th August 2023 Text data language: Bahasa Melayu (informal) Contributed to: https://github.com/huseinzol05/malaysian-dataset Pull request: https://github.com/huseinzol05/malaysian-dataset/pull/245 textn<1K0 likes36 downloads3y agoHugging Face14ammarxix /crawl-mufti-negeri-sembilan Details Source: https://muftins.gov.my/ Scrap date: 26/08/2023 textn<1K0 likes36 downloads3y agoHugging Face15crawlora-net /tiktok-engagement-index TikTok Engagement Index — engagement rate by niche An open dataset of TikTok engagement rates across 13 content niches, measured from 4,509 public videos (each with 10,000+ views) in 2026. Headline: Education content engages hardest (16.5% mean engagement rate); Tech trails at 5.9% — nearly 3× lower. The median video across all niches sits at 10.6%. Engagement rate = (likes + comments + shares + saves) / views, computed per video. 📊 Interactive study & chart:… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/tiktok-engagement-index.tabular1K<n<10K1 likes36 downloads2mo agoHugging Face16wanadzhar913 /crawl-mat-gaming TLDR Website: mat-gaming Num. pages scraped: 49 Remaining pages: 0 Date of scraping: 4th August 2023 Text data language: Bahasa Melayu Contributed to: https://github.com/huseinzol05/malaysian-dataset Pull request: https://github.com/huseinzol05/malaysian-dataset/pull/242 textn<1K0 likes31 downloads3y agoHugging Face17syafie-nzm /crawl-tamil.goodreturns.inscraped from https://tamil.goodreturns.in/topic/malaysia textn<1K0 likes29 downloads3y agoHugging Face18aisyahhrazak /crawl-medmalaytext10K<n<100K0 likes29 downloads3y agoHugging Face19crawlfeeds /walmart-reviews-dataset 🛒 Walmart Product Reviews Dataset (6.7K Records) This dataset contains 6,700+ structured customer reviews from Walmart.com. Each entry includes product-level metadata along with review details, making it ideal for small-scale machine learning models, sentiment analysis, and ecommerce insights. 📑 Dataset Fields Column Description url Direct product page URL name Product name/title sku Product SKU (Stock Keeping Unit) price Product price (numeric, USD)… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/walmart-reviews-dataset.tabulartext-classification1K<n<10K0 likes29 downloads1y agoHugging Face20pixeloffice /ai-crawlers-ip-and-user-agent-registry Known AI Crawlers, User-Agents & Verified IP Ranges Registry (2026) Official aggregated registry of all major AI Scrapers, Search Bots, and LLM Crawlers operating on the web. It includes verified User-Agent strings, policy block targets (for robots.txt), and official CIDR IP ranges for firewall configuration. Published by Pixel Office EU. Purpose & Utility As AI web indexing scales exponentially, website administrators and system engineers face the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/ai-crawlers-ip-and-user-agent-registry.texttext-classificationn<1K0 likes29 downloads1mo agoHugging Face21wanadzhar913 /crawl-bikesrepublic TLDR website: bikesrepublic num. of webpages scraped: 6,969 link to dataset: https://huggingface.co/datasets/wanadzhar913/crawl-bikesrepublic last date of scraping: 10th September 2023 status: complete pull request: https://github.com/huseinzol05/malaysian-dataset/pull/291 contributed to: https://github.com/huseinzol05/malaysian-dataset text1K<n<10K0 likes28 downloads3y agoHugging Face22aisyahhrazak /crawl-malaysiagazetteAbout Data scraped from https://malaysiagazette.com/ on 4.7.2023 text100K<n<1M0 likes27 downloads3y agoHugging Face23wanadzhar913 /crawl-timchewTLDR website: timchew num. of webpages scraped: 839 link to dataset: https://huggingface.co/datasets/wanadzhar913/crawl-timchew last date of scraping: 10th September 2023 status: complete pull request: https://github.com/huseinzol05/malaysian-dataset/pull/313 contributed to: https://github.com/huseinzol05/malaysian-dataset textn<1K0 likes27 downloads3y agoHugging Face24wanadzhar913 /crawl-techrakyat website: techrakyat num. of webpages scraped: 220 contributed to: https://github.com/huseinzol05/malaysian-dataset textn<1K0 likes26 downloads3y agoHugging Face25philschmid /crawl-datasettextn<1K0 likes25 downloads3y agoHugging Face26ammarxix /crawl-mufti-pahang Details Source: https://mufti.pahang.gov.my/ Scrap date: 26/08/2023 textn<1K0 likes25 downloads3y agoHugging Face27wanadzhar913 /crawl-vikatan-my TLDR website: Vikatan-MY num. of webpages scraped: 65 (7 locked behind paywall) link to dataset: https://huggingface.co/datasets/wanadzhar913/crawl-vikatan-my/resolve/main/vikatan-my-scraped-data.jsonl date of scraping: 21st October 2023 pull request: mesolitica/malaysian-dataset#353 contributed to: https://github.com/mesolitica/malaysian-dataset textn<1K0 likes25 downloads3y agoHugging Face28aisyahhrazak /crawl-agendadailyhttps://www.agendadaily.com/ text10K<n<100K0 likes23 downloads3y agoHugging Face29aisyahhrazak /crawl-fliphtmlFliphtml pdf text version Search Query: Melayu text1K<n<10K0 likes23 downloads3y agoHugging Face30ammarxix /crawl-doktorbudaktextn<1K0 likes22 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.