datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Web_Scraper_Datakorean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.legal_scrapers
JusticeDAO Legal Scrapers
Research collectors that download official legislative sources only
(national gazettes and official open-data portals). Intended for building
reproducible legal-text corpora, not for production legal research products.
Not legal advice. The official gazette of each jurisdiction prevails.
License
AGPL-3.0 for the collector scripts in this repository.
Contents
scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.finra-brokercheck-scraper
FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures
Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows.
Rows in this dataset
1,430
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.neptun.scraper
Data in this dataset
Docker & NPM
Scraped using crawl4ai.
The NPM and Docker data was scraped from docs.docker.com and docs.npmjs.com and processed using GPT-4 resulting in docker_documentation.jsonl and npm_documentation.jsonl.
The file training-data-v1.jsonl also includes Titanium, dockerNLcommands and docker_ps.
GitHub
Scraped using firecrawl.
The GitHub data was scraped from docs.github.com/en using firecrawl.A few pages might be missing in the… See the full description on the dataset page: https://huggingface.co/datasets/neptun-org/neptun.scraper.greenhouse-jobs-scraper
Greenhouse Jobs Scraper
Scrape every public job posting from any Greenhouse company job board: title, department, location, remote flag, seniority, advertised salary, full description and apply URL.
Rows in this dataset
14,091
Fields
36
Collector runs behind it
61
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/greenhouse-jobs-scraper/ — 92 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/greenhouse-jobs-scraper.grants-gov-scraper
Grants.gov Scraper · Grant Opportunities, Agencies & Awards
Scrape US federal grant opportunities, funding announcements, and agency award notices from Grants.gov by keyword, agency, category, eligibility, and status.
Rows in this dataset
1,488
Fields
12
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/grants-gov-scraper.shopify-store-products-scraper
Shopify Store Scraper
Scrape products, prices, discounts, variants and stock from any Shopify store's public product JSON. No login, no API key, no headless browser.
Rows in this dataset
11,720
Fields
33
Collector runs behind it
72
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-store-products-scraper/ — 1,271 entity pages
Run the collector yourself
https://apify.com/reapx/shopify-store-products-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-store-products-scraper.boardgamegeek-scraper
BoardGameGeek Scraper · Games, Ratings, Designers & Mechanics
Scrape board games, release years, player counts, categories, mechanics, designers, artists, and publishers from BoardGameGeek. HTTP only, pay-per-event pricing.
Rows in this dataset
437
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/boardgamegeek-scraper.shopify-app-store-scraper
Shopify App Store Scraper
List Shopify App Store apps and get one structured row per app, with developer, star rating, review count, every advertised pricing plan and the app's rank in its category.
Rows in this dataset
3,880
Fields
26
Collector runs behind it
87
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/shopify-app-store-scraper/ — 1,625 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/shopify-app-store-scraper.stackoverflow-scraper
StackOverflow Scraper
Scrape Stack Overflow questions, answers, tags and user profiles through the public Stack Exchange API. Filter by tag, score, date, accepted status and full-text search. No login, no browser.
Rows in this dataset
16,719
Fields
47
Collector runs behind it
57
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/stackoverflow-scraper/ — 9,841 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/stackoverflow-scraper.sec-edgar-scraper
SEC EDGAR Scraper
Search SEC EDGAR by ticker, CIK, SIC industry code, or full-text phrase and get one structured row per filing, with company classification, period dates, direct document links and optional XBRL financials.
Rows in this dataset
2,691
Fields
27
Collector runs behind it
64
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/sec-edgar-scraper/ — 771 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/sec-edgar-scraper.app-store-reviews-scraper
App Store Reviews Scraper
Scrape Apple App Store reviews, star ratings and app version history for any iOS app in any country storefront. No login, no API key.
Rows in this dataset
23,048
Fields
43
Collector runs behind it
88
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/app-store-reviews-scraper/ — 176 entity pages
Run the collector yourself
https://apify.com/reapx/app-store-reviews-scraper
What this is… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/app-store-reviews-scraper.nvd-cve-scraper
NVD CVE Scraper · Vulnerabilities, CVSS Scores, Vendors & CWEs
Scrape National Vulnerability Database (NVD) CVE records, CVSS v2/v3/v4 severity scores, CWE weakness classifications, vendor products, and exploit references.
Rows in this dataset
1,403
Fields
19
Collector runs behind it
36
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/nvd-cve-scraper.kalshi-scraper
Kalshi Scraper · Event Contracts, Markets, Prices & Volume
Scrape Kalshi prediction markets, event contracts, option pricing, order book quotes, trading volume, open interest, and resolution rules. Export structured JSON, CSV, or Excel data.
Rows in this dataset
7,649
Fields
34
Collector runs behind it
49
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/kalshi-scraper.arxiv-papers-scraper
arXiv Papers Scraper
Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window.
Rows in this dataset
21,722
Fields
26
Collector runs behind it
92
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.doaj-scraper
DOAJ Scraper · Open Access Journals, Articles & Publishers
Scrape open access research articles, DOIs, authors, subjects, publishers, and abstracts from the Directory of Open Access Journals (DOAJ) API. Features pay-per-event pricing and automatic backoff.
Rows in this dataset
1,720
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/doaj-scraper.github-repo-scraper
GitHub Repo Scraper · Repositories, Stars, Topics & Languages
Scrape GitHub repositories by language, topic, star count, license, organization, and pushed date window. Returns clean structured repo metrics and metadata without authentication.
Rows in this dataset
2,481
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/github-repo-scraper/ — 2,071 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/github-repo-scraper.remoteok-jobs-scraper
RemoteOK Jobs Scraper
Search RemoteOK and get one structured row per remote job, with company, role, tags, posted date, published salary range where given, and a direct apply link.
Rows in this dataset
3,001
Fields
30
Collector runs behind it
53
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/remoteok-jobs-scraper/ — 1,108 entity pages
Run the collector yourself
https://apify.com/reapx/remoteok-jobs-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/remoteok-jobs-scraper.lever-jobs-scraper
Lever Jobs Scraper · Job Postings, Teams, Locations & Salary
Scrape open job postings, departments, locations, remote status, compensation and application URLs from any Lever company job board.
Rows in this dataset
8,716
Fields
20
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a keyword list: a row exists… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/lever-jobs-scraper.semantic-scholar-scraper
Semantic Scholar Scraper · Papers, Authors, Citations & Venues
Scrape academic research papers, authors, citations, venues, and open-access metadata from Semantic Scholar API. Features rate-limit backoff resilience and pay-per-event pricing.
Rows in this dataset
450
Fields
19
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/semantic-scholar-scraper.wikipedia-scraper
Wikipedia Scraper · Articles, Extracts, Categories & Links
Extract Wikipedia articles, full text, lead summaries, categories, internal links, page views, and metadata across languages. HTTP-only Wikipedia API scraper for research, LLM datasets, and knowledge graphs.
Rows in this dataset
2,500
Fields
11
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wikipedia-scraper.remote-jobs-scraper
Remote Jobs Scraper · Remote Openings, Salary, Tags & Company
Scrape open remote jobs from Remotive API with company name, job title, salary, category, tags, locations, and direct apply link. Emits companyName per row.
Rows in this dataset
1,179
Fields
18
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/remote-jobs-scraper.wordpress-plugins-scraper
WordPress Plugins Scraper · Plugins, Installs & Ratings
Scrape WordPress plugins directory by tags, search queries, author accounts, install bands, and rating filters. Extract ratings, active installs, tags, release details, and author links.
Rows in this dataset
1,660
Fields
23
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/wordpress-plugins-scraper.pubmed-scraper
PubMed Scraper · Papers, Authors, Journals & MeSH Terms
Scrape academic research papers, authors, journals, MeSH terms, abstracts, and open-access metadata from NCBI PubMed API. Features HTTP backoff resilience and pay-per-event pricing.
Rows in this dataset
1,847
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/pubmed-scraper.openalex-scraper
OpenAlex Scraper · Works, Authors, Institutions & Citations
Scrape scholarly works, papers, citations, authors, institutions, and open-access metadata from the OpenAlex API. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
2,387
Fields
18
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/openalex-scraper/ — 2,387 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/openalex-scraper.zenodo-scraper
Zenodo Scraper · Research Records, DOIs, Authors & Files
Scrape open research records, DOIs, publications, datasets, software, authors, and file metadata from Zenodo. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
1,965
Fields
27
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/zenodo-scraper.crossref-scraper
Crossref Scraper · DOI Metadata, Authors, Journals & Citations
Scrape scholarly DOI metadata, works, journal articles, authors, citations, funding, and licenses from the Crossref REST API. Fast HTTP scraper with pay-per-event pricing.
Rows in this dataset
2,492
Fields
29
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/crossref-scraper/ — 2,492 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/crossref-scraper.steam-reviews-scraper
Steam Reviews Scraper · Game Reviews, Ratings & Playtime
Scrape Steam game reviews, ratings, playtime, helpfulness votes, and purchase types across any Steam App ID. Fast HTTP scraper, no login required.
Rows in this dataset
9,260
Fields
26
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a keyword list: a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/steam-reviews-scraper.apple-podcasts-scraper
Apple Podcasts Scraper · Shows, Episodes, Genres & Rankings
Scrape Apple Podcasts catalog, shows, episodes, top charts, genres, and rankings. HTTP-only iTunes Search API scraper for audio analytics, podcast discovery, and media datasets.
Rows in this dataset
2,189
Fields
20
Collector runs behind it
51
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/apple-podcasts-scraper.
