Scraper
Datasets
All datasets matching “Scraper”Web_Scraper_Datakorean_data_scraper_aihub
Korean Data Scraper — AI-Hub Local Corpus
상태: 비공개 (private) 저장소입니다.
Korean_data_scraper 프로젝트의 aihub_local 소스가 생성한 코퍼스입니다. 국립정보화진흥원 AI-Hub(aihub.or.kr)에서 내려받은 여러 데이터셋의 로컬 zip 압축 파일을 압축 해제하고, 그 안의 json 파일에서 본문 텍스트를 추출한 결과입니다. 샘플로 확인한 내용 중에는 뉴스 기사(신문기사) 카테고리의 AI-Hub 데이터셋에서 추출된 텍스트가 포함되어 있습니다.
스키마
파일당 1개 레코드(JSONL)이며, 다음과 같은 필드를 가집니다.
{"id": "aihub_local:<zip 파일명>:<json 파일명>", "source": "aihub_local", "text": "...", "url": null, "license": "per-dataset -- check the… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/korean_data_scraper_aihub.mastodon-scraper
Mastodon Scraper
Read public Mastodon posts by hashtag, instance timeline, trending or account, and get one structured row per post with the author, engagement counts, hashtags, media and outbound links.
Rows in this dataset
24,226
Fields
67
Collector runs behind it
66
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/mastodon-scraper/ — 7,898 entity pages
Run the collector yourself
https://apify.com/reapx/mastodon-scraper… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/mastodon-scraper.legal_scrapers
JusticeDAO Legal Scrapers
Research collectors that download official legislative sources only
(national gazettes and official open-data portals). Intended for building
reproducible legal-text corpora, not for production legal research products.
Not legal advice. The official gazette of each jurisdiction prevails.
License
AGPL-3.0 for the collector scripts in this repository.
Contents
scrapers/ includes country collectors, EU/EUR-Lex, shared helpers… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/legal_scrapers.wikimedia_scrapergleif-lei-scraper-sample-data
GLEIF LEI Scraper
Scrape global legal entities from the official GLEIF LEI database — no login, no API key, no blocking. 3.3M+ entities worldwide with legal name, address, jurisdiction, legal form, status and registration data. Filter by country and status. Tens of thousands per run.
What the actor scrapes
🏛️ GLEIF LEI Scraper — Global Legal Entity Identifier Data to JSON/CSV/Excel Scrape global legal entities straight from the official GLEIF API — the worldwide… See the full description on the dataset page: https://huggingface.co/datasets/logiover/gleif-lei-scraper-sample-data.
