datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icelandic-common-crawl-corpus-IC3This is the Icelandic Common Crawl Corpus (IC3).
crawl-karangan-net-komsasagillm-crawl-datagoddess-crawlcrawl-lirik-lagu-dot-netlink_crawl_beritacatalan_government_crawling
Dataset Card for Catalan Government Crawling
Dataset Summary
The Catalan Government Crawling Corpus is a 39-million-token web corpus of Catalan built from the web. It has been obtained by crawling the .gencat domain and subdomains, belonging to the Catalan Government during September and October 2020. It consists of 39,117,909 tokens, 1,565,433 sentences and 71,043 documents. Documents are separated by single new lines. It is a subcorpus of the Catalan Textual Corpus.… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalan_government_crawling.CRAwLeR-PL
CRAwLeR-PL — Cross-Reference Aware Legal Retrieval (Polish)
CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Polish instance, built on Polish legal Acts obtained through the ELI API.
It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-PL.vi_math_problem_crawl
Dataset Card for Vietnamese Elementary Math Knowledge and Workbook
Dataset Summary
The data includes information about elementary school math knowledge in Vietnam, as well as exercises compiled from books. This is a crawlable dataset that can be trained for text generation tasks.
Supported Tasks and Leaderboards
Languages
The majority of the data is in Vietnamese, but there is still some English from some bilingual workbooks.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hllj/vi_math_problem_crawl.CRAwLeR-DK
CRAwLeR-DK — Cross-Reference Aware Legal Retrieval (Danish)
CRAwLeR measures context-aware (contextual) chunk retrieval: specifically, legal cross-reference retrieval. To rank the right chunk, a retriever must use information from other chunks that the target chunk cross-references. This is the Danish instance, built on Danish legal documents from Retsinformation.
It is the first dataset for context-aware chunk retrieval to carefully consider construct validity and inspect… See the full description on the dataset page: https://huggingface.co/datasets/macja/CRAwLeR-DK.ai-crawler-user-agents
AI Crawler User Agents
Machine-readable list of all 28 known AI crawler and agent user-agent strings —
GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, Bytespider,
Applebot-Extended, and more — with each bot's operator, purpose, observed
robots.txt compliance, and documented crawl-delay support.
Fields
Field
Description
userAgent
Token to match in robots.txt / server logs (e.g. GPTBot)
operator
Company running the crawler
purpose… See the full description on the dataset page: https://huggingface.co/datasets/osamamumtaz01/ai-crawler-user-agents.crawl-malaysian-websitecrawl-leaazleeya
TLDR
Website: leaazleeya
Num. of webpages: 543
Num. of webpages scraped: 543
Num. articles successfully extracted: 534
Remaing webpages to be scraped: 0
Scraped on: 5th August 2023
Text data language: Bahasa Melayu (informal)
Contributed to: https://github.com/huseinzol05/malaysian-dataset
Pull request: https://github.com/huseinzol05/malaysian-dataset/pull/245
crawl-mufti-negeri-sembilan
Details
Source: https://muftins.gov.my/
Scrap date: 26/08/2023
tiktok-engagement-index
TikTok Engagement Index — engagement rate by niche
An open dataset of TikTok engagement rates across 13 content niches, measured from 4,509 public videos (each with 10,000+ views) in 2026.
Headline: Education content engages hardest (16.5% mean engagement rate); Tech trails at 5.9% — nearly 3× lower. The median video across all niches sits at 10.6%.
Engagement rate = (likes + comments + shares + saves) / views, computed per video.
📊 Interactive study & chart:… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/tiktok-engagement-index.crawl-mat-gaming
TLDR
Website: mat-gaming
Num. pages scraped: 49
Remaining pages: 0
Date of scraping: 4th August 2023
Text data language: Bahasa Melayu
Contributed to: https://github.com/huseinzol05/malaysian-dataset
Pull request: https://github.com/huseinzol05/malaysian-dataset/pull/242
crawl-tamil.goodreturns.inscraped from https://tamil.goodreturns.in/topic/malaysia
crawl-medmalaywalmart-reviews-dataset
🛒 Walmart Product Reviews Dataset (6.7K Records)
This dataset contains 6,700+ structured customer reviews from Walmart.com. Each entry includes product-level metadata along with review details, making it ideal for small-scale machine learning models, sentiment analysis, and ecommerce insights.
📑 Dataset Fields
Column
Description
url
Direct product page URL
name
Product name/title
sku
Product SKU (Stock Keeping Unit)
price
Product price (numeric, USD)… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/walmart-reviews-dataset.ai-crawlers-ip-and-user-agent-registry
Known AI Crawlers, User-Agents & Verified IP Ranges Registry (2026)
Official aggregated registry of all major AI Scrapers, Search Bots, and LLM Crawlers operating on the web. It includes verified User-Agent strings, policy block targets (for robots.txt), and official CIDR IP ranges for firewall configuration.
Published by Pixel Office EU.
Purpose & Utility
As AI web indexing scales exponentially, website administrators and system engineers face the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/pixeloffice/ai-crawlers-ip-and-user-agent-registry.crawl-bikesrepublic
TLDR
website: bikesrepublic
num. of webpages scraped: 6,969
link to dataset: https://huggingface.co/datasets/wanadzhar913/crawl-bikesrepublic
last date of scraping: 10th September 2023
status: complete
pull request: https://github.com/huseinzol05/malaysian-dataset/pull/291
contributed to: https://github.com/huseinzol05/malaysian-dataset
crawl-malaysiagazetteAbout
Data scraped from https://malaysiagazette.com/
on 4.7.2023
crawl-timchewTLDR
website: timchew
num. of webpages scraped: 839
link to dataset: https://huggingface.co/datasets/wanadzhar913/crawl-timchew
last date of scraping: 10th September 2023
status: complete
pull request: https://github.com/huseinzol05/malaysian-dataset/pull/313
contributed to: https://github.com/huseinzol05/malaysian-dataset
crawl-techrakyat
website: techrakyat
num. of webpages scraped: 220
contributed to: https://github.com/huseinzol05/malaysian-dataset
crawl-datasetcrawl-mufti-pahang
Details
Source: https://mufti.pahang.gov.my/
Scrap date: 26/08/2023
crawl-vikatan-my
TLDR
website: Vikatan-MY
num. of webpages scraped: 65 (7 locked behind paywall)
link to dataset: https://huggingface.co/datasets/wanadzhar913/crawl-vikatan-my/resolve/main/vikatan-my-scraped-data.jsonl
date of scraping: 21st October 2023
pull request: mesolitica/malaysian-dataset#353
contributed to: https://github.com/mesolitica/malaysian-dataset
crawl-agendadailyhttps://www.agendadaily.com/
crawl-fliphtmlFliphtml pdf text version
Search Query:
Melayu
crawl-doktorbudak
