CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hudsongouge /AoPS-Scrape AoPS-Scrape Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints. Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session. Splits Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split: Split Rows Notes deduplicated 29,964 One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tabularquestion-answering10K<n<100K2 likes4.4k downloads2mo agoHugging Face02Gugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.8k downloads19d agoHugging Face03mkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes353 downloads24d agoHugging Face04scrapegraphai /scrapegraphai-100k ScrapeGraphAI-100k Dataset Summary ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.tabulartext-generation10K<n<100K26 likes160 downloads2mo agoHugging Face05scrapegraphai /scrapegraph-100k-finetuning ScrapeGraphAI 100k finetuning Dataset Summary A finetuning-ready derivative of ScrapeGraphAI-100k: schema-constrained web extraction examples where a model must produce JSON conforming to a user-defined JSON schema given Markdown-converted page content. Split Rows Targets train 25,244 GPT-5-nano regenerated targets test 2,808 GPT-5-nano regenerated targets human_eval 100 Human labeled extractions (evaluation only) Important: train/test… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraph-100k-finetuning.textfeature-extraction10K<n<100K2 likes88 downloads2mo agoHugging Face06llmtraining-scraper /discord-messages Discord Messages Dataset Description This dataset contains 6.2 million anonymized messages extracted from public Discord servers. All personal identifying information (user IDs, server IDs, channel IDs, timestamps) has been removed. Only the raw message text remains. The data is formatted as plain text with one message per line, making it ideal for: Language model pre-training Fine-tuning chatbots Sentiment analysis Toxicity detection Slang and language evolution… See the full description on the dataset page: https://huggingface.co/datasets/llmtraining-scraper/discord-messages.texttext-generation1M<n<10M0 likes61 downloads1mo agoHugging Face07sayurio /ekpatagolpo-scrape-bangla-literature Ekpatagolpo Bengali Stories Archive Request More ScrapesOrder Private Scrapes Overview This repository contains a large-scale, curated text dataset scraped from ekpatagolpo.com. The primary goal of this archive is to preserve a massive collection of purely human-written Bengali literature and stories (Bangla Golpo), creating a distinct record of human creativity separate from AI-generated text. Purpose and Usage This dataset is published publicly under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/ekpatagolpo-scrape-bangla-literature.texttext-generation10K<n<100K1 likes59 downloads6mo agoHugging Face08mrcuddle /Nifty-Authoritarian-ScrapeData Scrape from LGBT Literature Archive Nifty.Org -Category: Authoritarian texttext-classification1K<n<10K1 likes41 downloads2y agoHugging Face09thibble /paper2env-scraped Paper2Env — scraped arXiv One row per task. Each task is a paper-reproduction subtask with a verification script (verify.sh) that scores submissions, plus a text-only git diff patch against an upstream GitHub repo at a pinned commit. Per-task binary artefacts (paper PDF, assets, expected outputs for grading, binary file additions to the student repo) live in the companion repo thibble/paper2env-artifacts under scraped/<paper_id>/<task_id>.tar.gz. Reconstruct a task… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-scraped.texttext-generationn<1K0 likes32 downloads5mo agoHugging Face10sayurio /jugantor.com-scrape-bangla Jugantor News Archive (Bangla) Overview This repository contains a comprehensive text dataset scraped from jugantor.com, one of the leading Bengali daily newspapers in Bangladesh. The primary goal of this archive is to preserve a massive collection of purely human-written journalism, editorials, and news reports, creating a distinct record of human-authored text separate from AI-generated content. Purpose and Usage This dataset is published publicly and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/jugantor.com-scrape-bangla.imagetext-generation10K<n<100K1 likes31 downloads6mo agoHugging Face11sayurio /startech.com.bd-web-scrape Star Tech Product Archive Overview This repository contains a comprehensive dataset scraped from startech.com.bd, a major retailer of computers, hardware components, and consumer electronics in Bangladesh. The dataset serves as a structured archive of product catalogs, technical specifications, pricing, and descriptions. Purpose and Usage This dataset is published publicly and strictly for educational, research, and analytical purposes. It is an excellent… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/startech.com.bd-web-scrape.texttext-generation10K<n<100K1 likes27 downloads6mo agoHugging Face12sayurio /dainikparibarton-web-scrape-bangla Dainik Paribarton News Archive (Bangla) Request More ScrapesOrder Private Scrapes Overview This repository contains a text dataset scraped from dainikparibarton.com, a Bengali online news portal covering national, regional, sports, and political news in Bangladesh. The primary goal of this archive is to preserve a collection of purely human-written journalism and regional reporting, creating a distinct record of human-authored text separate from AI-generated content.… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/dainikparibarton-web-scrape-bangla.imagetext-generation10K<n<100K1 likes24 downloads6mo agoHugging Face13mkd-minju /korean_data_scraper_kakaotalk Korean Data Scraper — KakaoTalk Export Corpus © 주식회사 MKD(MKD Inc.) — 모든 권리 보유 (All rights reserved). 본 데이터셋은 주식회사 MKD(MKD Inc.)가 자체적으로 생성·소유한 내부 데이터이며, 외부 공개 라이선스가 부여되지 않은 사내 전용(proprietary) 자료입니다. Korean_data_scraper 프로젝트의 kakaotalk_export 소스가 생성한 코퍼스입니다. 사용자 본인이 참여자인 KakaoTalk 그룹 채팅방 2곳의 공식 대화 내보내기(대화 내보내기) 파일을 파싱한 결과이며, 웹 스크래핑이 아니라 사용자가 직접 소유한 원본 데이터를 가공한 것입니다. 비공개(Private) 저장소입니다. 원본 채팅에 사내 인프라 정보와 다수 참여자의 실명이 포함되어 있어, 아래 정제 과정을 거쳤더라도 외부 공개를 의도한 데이터셋이 아닙니다.… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_kakaotalk.texttext-generationn<1K0 likes19 downloads24d agoHugging Face14benxh /qysh-me-shqip-scrape-datasetA small/complete web scrape of the old qysh.me website now hosted on dua.com. texttext-generationn<1K2 likes17 downloads3y agoHugging Face15sayurio /storymirror.com-web-scrape StoryMirror Multilingual Literature Archive Overview This repository contains a large-scale text dataset scraped from storymirror.com, a prominent digital platform for Indian literature. The primary goal of this archive is to preserve a massive, multilingual collection of purely human-written stories, poems, and quotes, creating a distinct record of human creativity and storytelling across various Indian languages. Purpose and Usage This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/storymirror.com-web-scrape.imagetext-generation10K<n<100K1 likes15 downloads6mo agoHugging Face16sayurio /bangla-kobita-scrape-bangla-literature Bangla Kobita Poetry Archive Overview This repository contains a curated text dataset of Bengali poetry scraped from the web, primarily targeting comprehensive poetry platforms like bangla-kobita.com. The primary goal of this archive is to preserve a rich collection of purely human-written Bengali poems (Bangla Kobita), creating a distinct record of human artistic expression, emotion, and linguistic rhythm separate from AI-generated text. Purpose and… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-kobita-scrape-bangla-literature.imagetext-generation100K<n<1M1 likes12 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.