datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GIT-SCRAPED
🚀 SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
📂 Repository Structure
All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.apebench-scraped
APEBench Scraped
A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0.
⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.IMDB-PUBLIC-SCRAPED
Hello World with Hugging Face
Current Date: 2025-03-19 04:36:42.698271
So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later.
It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks.
I'll be working out the problems and getting the scraper working correctly at some point soon.
english-manhwa-scrape
English Manhwa Scrape Dataset
Request More ScrapesOrder Private Scrapes
Note: The total files are about 150+GB and my internet connection is slow af. So I'll be uploading them in batches
File Structure:
files/
- Manhwa 1 Name
- Chapter 1.cbz
- Chapter 2.cbz
- Chapter 3.cbz
- ......
- Chapter n.cbz
- Manhwa 2 Name
- Chapter 1.cbz
- Chapter 2.cbz
- Chapter 3.cbz
- ......
- Chapter n.cbz
Due to maximum 10000 files in a repo limit of huggingface, I had to further… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/english-manhwa-scrape.github_repo_scrapedScraped python packages from github.
apebench-scraped-old
APEBench-scraped (old)
All datasets scraped from the APEBench benchmark suite with the version used
for the Neurips
submission.
Download
Download without large files
GIT_LFS_SKIP_SMUDGE=1 git clone git@hf.co:datasets/thuerey-group/apebench-scraped-old
Afterwards, you can inspect the repository and download the files you need. For
example, for 1d_diff_adv:
git lfs install
git lfs pull -I "data/1d_diff_adv*"
Alternatively, you can download the entire repository with large… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped-old.tatoeba-nusax-scrape-mt-concatWeb_Scraper_Datakorean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.mastodon-scraper
Mastodon Scraper
Read public Mastodon posts by hashtag, instance timeline, trending or account, and get one structured row per post with the author, engagement counts, hashtags, media and outbound links.
Rows in this dataset
24,226
Fields
67
Collector runs behind it
66
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/mastodon-scraper/ — 7,898 entity pages
Run the collector yourself
https://apify.com/reapx/mastodon-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/mastodon-scraper.legal_scrapers
JusticeDAO Legal Scrapers
Research collectors that download official legislative sources only
(national gazettes and official open-data portals). Intended for building
reproducible legal-text corpora, not for production legal research products.
Not legal advice. The official gazette of each jurisdiction prevails.
License
AGPL-3.0 for the collector scripts in this repository.
Contents
scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.wikimedia_scrapergleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.korean_data_scraper_wikipedia
Korean Data Scraper — Wikipedia Dump Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 wikipedia_dump 소스가 생성한 코퍼스입니다. 한국어 위키백과(ko.wikipedia.org) 덤프를 파싱하여 문서 본문 텍스트를 추출한 결과입니다.
스키마
파일당 1개 레코드(JSONL)이며, korean_data_scraper_kakaotalk와 동일한 공통 스키마를 따릅니다.
{"id": "kowiki:70773", "source": "wikipedia_dump", "text": "..."}
id는 kowiki:<문서 ID> 형태이며, text는 해당 위키백과 문서의 본문 텍스트입니다.
데이터 규모 및 한계
전체 32개 샤드(shard-00000~00031)로 구성되며, 총 용량은 약 2.24GB입니다. 이 중… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_wikipedia.Darkforums-Scraped-Dump-osint-Datapenetration_testing_scraped_dataset
Dataset Card for "penetration_testing_scraped_dataset"
More Information needed
linkedin-job-scrape
LinkedIn DS/ML Job Postings
Daily snapshots of data science / machine learning job postings scraped from LinkedIn, one parquet file per scrape run. All splits share one schema (2023 → today); historical splits were migrated in 2026-07 — the original 8-column data is preserved at revision v1-schema.
The same job_id recurs across splits (a posting stays live for days) — that's the longitudinal signal. For a unique-jobs view, dedup on job_id keeping the row with max scrape_dt.… See the full description on the dataset page: https://huggingface.co/datasets/ryang2/linkedin-job-scrape.tool-scraper_daniel_20260827_124546This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"left_carriage_joint.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-scraper_daniel_20260827_124546.vi.wikipedia-scraped-dataSorry for not having English.
bru
fitcheck-scraped-multiviewtelegram-scraped-datafitcheck-scraped-v1scrapegraphai-100k
ScrapeGraphAI-100k
Dataset Summary
ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.DrugHub-scrape
DrugHub Market Snapshot, September 2026
A complete, text-only capture of the public listing, vendor, and review pages of
DrugHub, a Monero-only darknet market operating since 2023. Everything here was
visible to any visitor without an account. Doesn't include any images.
Collected 16-17 September 2026. Enriched with model-derived labels
(typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment
table and the Enrichment section below.
What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.English_French_Webpages_Scraped_Translated
English French Webpages Scraped Translated
Dataset Summary
French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.scraped-chatgpt-conversations
Dataset Card for Dataset Name
Dataset Summary
scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.bangla-site-scrape
Bangladeshi Web Scraped Dataset
Request More ScrapesOrder Private Scrapes
Dataset Description
This dataset is a comprehensive collection of scraped web data from various Bangladeshi websites. It is designed to facilitate natural language processing (NLP) tasks for the Bengali language, including analysis of e-commerce trends, news classification, and general text generation.
The data is provided in JSONL (JSON Lines) format, making it easy to parse and integrate into… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-site-scrape.Garuda_Scrape
