CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SKT-NRS /GIT-SCRAPED 🚀 SKT-NRS / GIT-SCRAPED This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets. Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures. 📂 Repository Structure All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.1K<n<10K1 likes4.4k downloads3mo agoHugging Face02hudsongouge /AoPS-Scrape AoPS-Scrape Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints. Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session. Splits Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split: Split Rows Notes deduplicated 29,964 One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tabularquestion-answering10K<n<100K2 likes3.8k downloads2mo agoHugging Face03Gugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.8k downloads17d agoHugging Face04thuerey-group /apebench-scraped APEBench Scraped A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0. ⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.1 likes1.2k downloads2y agoHugging Face05AbstractPhil /IMDB-PUBLIC-SCRAPED Hello World with Hugging Face Current Date: 2025-03-19 04:36:42.698271 So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later. It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks. I'll be working out the problems and getting the scraper working correctly at some point soon. 1 likes903 downloads4mo agoHugging Face06sayurio /english-manhwa-scrape English Manhwa Scrape Dataset Request More ScrapesOrder Private Scrapes Note: The total files are about 150+GB and my internet connection is slow af. So I'll be uploading them in batches File Structure: files/ - Manhwa 1 Name - Chapter 1.cbz - Chapter 2.cbz - Chapter 3.cbz - ...... - Chapter n.cbz - Manhwa 2 Name - Chapter 1.cbz - Chapter 2.cbz - Chapter 3.cbz - ...... - Chapter n.cbz Due to maximum 10000 files in a repo limit of huggingface, I had to further… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/english-manhwa-scrape.image-to-text100K<n<1M2 likes808 downloads6mo agoHugging Face07Asib27 /github_repo_scrapedScraped python packages from github. 0 likes714 downloads2y agoHugging Face08thuerey-group /apebench-scraped-old APEBench-scraped (old) All datasets scraped from the APEBench benchmark suite with the version used for the Neurips submission. Download Download without large files GIT_LFS_SKIP_SMUDGE=1 git clone git@hf.co:datasets/thuerey-group/apebench-scraped-old Afterwards, you can inspect the repository and download the files you need. For example, for 1d_diff_adv: git lfs install git lfs pull -I "data/1d_diff_adv*" Alternatively, you can download the entire repository with large… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped-old.0 likes704 downloads2y agoHugging Face09thonyyy /tatoeba-nusax-scrape-mt-concattext100M<n<1B0 likes583 downloads2y agoHugging Face10Amin1600 /Web_Scraper_Datatext10K<n<100K1 likes570 downloads11h agoHugging Face11mkd-minju /korean_data_scraper_aihub Korean Data Scraper — AI-Hub Local Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다. 스키마 파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다. {"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.texttext-generation1M<n<10M0 likes352 downloads22d agoHugging Face12reapxdev /mastodon-scraper Mastodon Scraper Read public Mastodon posts by hashtag, instance timeline, trending or account, and get one structured row per post with the author, engagement counts, hashtags, media and outbound links. Rows in this dataset 24,226 Fields 67 Collector runs behind it 66 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/mastodon-scraper/ — 7,898 entity pages Run the collector yourself https://apify.com/reapx/mastodon-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/mastodon-scraper.10K<n<100K0 likes338 downloads2mo agoHugging Face13endomorphosis /legal_scrapers JusticeDAO Legal Scrapers Research collectors that download official legislative sources only (national gazettes and official open-data portals). Intended for building reproducible legal-text corpora, not for production legal research products. Not legal advice. The official gazette of each jurisdiction prevails. License AGPL-3.0 for the collector scripts in this repository. Contents scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.texttext-retrievaln<1K0 likes325 downloads13d agoHugging Face14JoannaCreatesArt /wikimedia_scraper0 likes318 downloads2y agoHugging Face15logiover /gleif-lei-scraper-sample-data GLEIF LEI Scraper Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run. What the actor scrapes 🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.textn<1K0 likes261 downloads4mo agoHugging Face16mkd-minju /korean_data_scraper_wikipedia Korean Data Scraper — Wikipedia Dump Corpus 상태: 비공개 (private) 저장소입니다. Korean_data_scraper 프로젝트의 wikipedia_dump 소스가 생성한 코퍼스입니다. 한국어 위키백과(ko.wikipedia.org) 덤프를 파싱하여 문서 본문 텍스트를 추출한 결과입니다. 스키마 파일당 1개 레코드(JSONL)이며, korean_data_scraper_kakaotalk와 동일한 공통 스키마를 따릅니다. {"id": "kowiki:70773", "source": "wikipedia_dump", "text": "..."} id는 kowiki:<문서 ID> 형태이며, text는 해당 위키백과 문서의 본문 텍스트입니다. 데이터 규모 및 한계 전체 32개 샤드(shard-00000~00031)로 구성되며, 총 용량은 약 2.24GB입니다. 이 중… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_wikipedia.text-generation1M<n<10M0 likes249 downloads22d agoHugging Face17DarkforumsHeiter /Darkforums-Scraped-Dump-osint-Datatextn<1K1 likes239 downloads3mo agoHugging Face18Isamu136 /penetration_testing_scraped_dataset Dataset Card for "penetration_testing_scraped_dataset" More Information needed text100K<n<1M14 likes238 downloads3y agoHugging Face19ryang2 /linkedin-job-scrape LinkedIn DS/ML Job Postings Daily snapshots of data science / machine learning job postings scraped from LinkedIn, one parquet file per scrape run. All splits share one schema (2023 → today); historical splits were migrated in 2026-07 — the original 8-column data is preserved at revision v1-schema. The same job_id recurs across splits (a posting stays live for days) — that's the longitudinal signal. For a unique-jobs view, dedup on job_id keeping the row with max scrape_dt.… See the full description on the dataset page: https://huggingface.co/datasets/ryang2/linkedin-job-scrape.text100K<n<1M0 likes232 downloads18d agoHugging Face20rbtrprjkt /tool-scraper_daniel_20260827_124546This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "joint_0.pos", "joint_1.pos", "joint_2.pos", "joint_3.pos", "joint_4.pos", "joint_5.pos", "left_carriage_joint.pos" ]… See the full description on the dataset page: https://huggingface.co/datasets/rbtrprjkt/tool-scraper_daniel_20260827_124546.tabularrobotics10K<n<100K0 likes222 downloads26d agoHugging Face21mondk /vi.wikipedia-scraped-dataSorry for not having English. bru 4 likes203 downloads1mo agoHugging Face22barryallen16 /fitcheck-scraped-multiviewimage100K<n<1M0 likes199 downloads8mo agoHugging Face23Sinfsr /telegram-scraped-data0 likes190 downloads7mo agoHugging Face24barryallen16 /fitcheck-scraped-v1image100K<n<1M0 likes181 downloads8mo agoHugging Face25scrapegraphai /scrapegraphai-100k ScrapeGraphAI-100k Dataset Summary ScrapeGraphAI-100k is a dataset of 93,695 real-world schema-constrained extraction events collected via opt-in telemetry of the ScrapeGraphAI open-source scraping library in Q2–Q3 2025. The dataset was derived from ~9 million raw PostHog telemetry events, deduplicated and balanced for schema diversity. Each example captures one LLM attempt to extract structured data from real web content under a user-defined JSON schema: the… See the full description on the dataset page: https://huggingface.co/datasets/scrapegraphai/scrapegraphai-100k.tabulartext-generation10K<n<100K26 likes149 downloads2mo agoHugging Face26trentmkelly /DrugHub-scrape DrugHub Market Snapshot, September 2026 A complete, text-only capture of the public listing, vendor, and review pages of DrugHub, a Monero-only darknet market operating since 2023. Everything here was visible to any visitor without an account. Doesn't include any images. Collected 16-17 September 2026. Enriched with model-derived labels (typesafe/jev-1.13) on 19 September 2026; see the listing_enrichment table and the Enrichment section below. What's in it… See the full description on the dataset page: https://huggingface.co/datasets/trentmkelly/DrugHub-scrape.tabular100K<n<1M3 likes137 downloads3d agoHugging Face27Nicolas-BZRD /English_French_Webpages_Scraped_Translated English French Webpages Scraped Translated Dataset Summary French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.texttranslation10M<n<100M3 likes117 downloads3y agoHugging Face28ar852 /scraped-chatgpt-conversations Dataset Card for Dataset Name Dataset Summary scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.question-answering100K<n<1M12 likes113 downloads3y agoHugging Face29sayurio /bangla-site-scrape Bangladeshi Web Scraped Dataset Request More ScrapesOrder Private Scrapes Dataset Description This dataset is a comprehensive collection of scraped web data from various Bangladeshi websites. It is designed to facilitate natural language processing (NLP) tasks for the Bengali language, including analysis of e-commerce trends, news classification, and general text generation. The data is provided in JSONL (JSON Lines) format, making it easy to parse and integrate into… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/bangla-site-scrape.text-generation10K<n<100K1 likes109 downloads6mo agoHugging Face30JSpergel /Garuda_Scrape0 likes107 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.