datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.arena.ai-code-leaderboard-scrapedlink: https://arena.ai/leaderboard/code/
scraped-web-datajurnal-malaysia-scraped
website: jurnal-malaysia
Number of pages scraped: 20
Number of posts scraped: 1938
Link to dataset on Huggingface
sfia-9-scraped
SFIA-9-Scraped Dataset
This repository contains the SFIA-9-Scraped dataset, a JSON collection of the Skills Framework for the Information Age (SFIA) version 9 categories and levels, scraped for non-commercial research use.
🚀 Dataset Overview
Name: SFIA-9-Scraped
Hugging Face: Programmer-RD-AI/sfia-9-scraped
DOI: 10.57967/hf/5746
Author: Ranuga Disansa Gamage
Revision: 89feeb8
Publisher: Hugging Face
Year: 2025
Use this dataset to build RAG systems, taxonomy-driven… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sfia-9-scraped.cypherhackz-scraped
website: cypherhackz
Number of pages scraped: 9
Number of posts scraped: 805
Link to dataset on Huggingface
scraped_notesdisease-scraped
disease-scraped
This dataset was uploaded automatically.
scraped_quizzesscraped-trends-database-236dafpi-scraped-datascrapedscraped_quizzesConnor-Data-Scraped-v1dvhs_scraped_data
