datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AoPS-Scrape
AoPS-Scrape
Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints.
Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session.
Splits
Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split:
Split
Rows
Notes
deduplicated
29,964
One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.korean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.scrapegraphai-100k
ScrapeGraphAI-100k
Dataset Summary
ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.scrapegraph-100k-finetuning
ScrapeGraphAI 100k finetuning
Dataset Summary
A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content.
Split
Rows
Targets
train
25,244
GPT-5-nano regenerated targets
test
2,808
GPT-5-nano regenerated targets
human_eval
100
Human labeled extractions (evaluation only)
Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.discord-messages
Discord Messages Dataset
Description
This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains.
The data is formatted as plain text with one message per line, making it ideal for:
Language model pre-training
Fine-tuning chatbots
Sentiment analysis
Toxicity detection
Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.ekpatagolpo-scrape-bangla-literature
Ekpatagolpo Bengali Stories Archive
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text.
Purpose and Usage
This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.Nifty-Authoritarian-ScrapeData Scrape from LGBT Literature Archive Nifty.Org
-Category: Authoritarian
paper2env-scraped
Paper2Env — scraped arXiv
One row per task. Each task is a paper-reproduction subtask with a verification
script (verify.sh) that scores submissions, plus a text-only git diff
patch against an upstream GitHub repo at a pinned commit.
Per-task binary artefacts (paper PDF, assets, expected outputs for grading,
binary file additions to the student repo) live in the companion repo
thibble/paper2env-artifacts under
scraped/<paper_id>/<task_id>.tar.gz.
Reconstruct a task… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-scraped.jugantor.com-scrape-bangla
Jugantor News Archive (Bangla)
Overview
This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content.
Purpose and Usage
This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.startech.com.bd-web-scrape
Star Tech Product Archive
Overview
This repository contains a comprehensive dataset scraped from startech.com.bd, a major retailer of computers, hardware components, and consumer electronics in Bangladesh. The dataset serves as a structured archive of product catalogs, technical specifications, pricing, and descriptions.
Purpose and Usage
This dataset is published publicly and strictly for educational, research, and analytical purposes. It is an excellent… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/startech.com.bd-web-scrape.dainikparibarton-web-scrape-bangla
Dainik Paribarton News Archive (Bangla)
Request More ScrapesOrder Private Scrapes
Overview
This repository contains a text dataset scraped from dainikparibarton.com, a Bengali online news portal covering national, regional, sports, and political news in Bangladesh. The primary goal of this archive is to preserve a collection of purely human-written journalism and regional reporting, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/dainikparibarton-web-scrape-bangla.korean_data_scraper_kakaotalk
Korean Data Scraper — KakaoTalk Export Corpus
© 주식회사 MKD(MKD Inc.) — 모든 권리 보유 (All rights reserved).
본 데이터셋은 주식회사 MKD(MKD Inc.)가 자체적으로 생성·소유한 내부 데이터이며,
외부 공개 라이선스가 부여되지 않은 사내 전용(proprietary) 자료입니다.
Korean_data_scraper 프로젝트의 kakaotalk_export 소스가 생성한 코퍼스입니다.
사용자 본인이 참여자인 KakaoTalk 그룹 채팅방 2곳의 공식 대화 내보내기(대화 내보내기)
파일을 파싱한 결과이며, 웹 스크래핑이 아니라 사용자가 직접 소유한 원본 데이터를
가공한 것입니다.
비공개(Private) 저장소입니다. 원본 채팅에 사내 인프라 정보와 다수 참여자의
실명이 포함되어 있어, 아래 정제 과정을 거쳤더라도 외부 공개를 의도한 데이터셋이
아닙니다.… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_kakaotalk.qysh-me-shqip-scrape-datasetA small/complete web scrape of the old qysh.me website now hosted on dua.com.
storymirror.com-web-scrape
StoryMirror Multilingual Literature Archive
Overview
This repository contains a large-scale text dataset scraped from storymirror.com, a prominent digital platform for Indian literature. The primary goal of this archive is to preserve a massive, multilingual collection of purely human-written stories, poems, and quotes, creating a distinct record of human creativity and storytelling across various Indian languages.
Purpose and Usage
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/storymirror.com-web-scrape.bangla-kobita-scrape-bangla-literature
Bangla Kobita Poetry Archive
Overview
This repository contains a curated text dataset of Bengali poetry scraped from the web, primarily targeting comprehensive poetry platforms like bangla-kobita.com. The primary goal of this archive is to preserve a rich collection of purely human-written Bengali poems (Bangla Kobita), creating a distinct record of human artistic expression, emotion, and linguistic rhythm separate from AI-generated text.
Purpose and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-kobita-scrape-bangla-literature.
