datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GIT-SCRAPED
🚀 SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
📂 Repository Structure
All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.apebench-scraped
APEBench Scraped
A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0.
⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.IMDB-PUBLIC-SCRAPED
Hello World with Hugging Face
Current Date: 2025-03-19 04:36:42.698271
So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later.
It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks.
I'll be working out the problems and getting the scraper working correctly at some point soon.
apebench-scraped-old
APEBench-scraped (old)
All datasets scraped from the APEBench benchmark suite with the version used
for the Neurips
submission.
Download
Download without large files
GIT_LFS_SKIP_SMUDGE=1 git clone git@hf.co:datasets/thuerey-group/apebench-scraped-old
Afterwards, you can inspect the repository and download the files you need. For
example, for 1d_diff_adv:
git lfs install
git lfs pull -I "data/1d_diff_adv*"
Alternatively, you can download the entire repository with large… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped-old.github_repo_scrapedScraped python packages from github.
vi.wikipedia-scraped-dataSorry for not having English.
bru
penetration_testing_scraped_dataset
Dataset Card for "penetration_testing_scraped_dataset"
More Information needed
telegram-scraped-datafitcheck-scraped-v1scraped-chatgpt-conversations
Dataset Card for Dataset Name
Dataset Summary
scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.English_French_Webpages_Scraped_Translated
English French Webpages Scraped Translated
Dataset Summary
French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.arena.ai-code-leaderboard-scrapedlink: https://arena.ai/leaderboard/code/
Darkforums-Scraped-Dump-osint-Datascraped-hotel-reviewsScrapedJobssobolo-corpus-multilang-scrapedhacker-news-scraped-storiesPolyMath-Scraped-Raw
PolyMath Scraped
PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems.
PolyMath addresses both issues through:
Data scraping: problems sourced from official competition PDFs absent from popular datasets, using a… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath-Scraped-Raw.scraped-episodesscraped-web-dataNYT-Connections-Verl-Scrapedjurnal-malaysia-scraped
website: jurnal-malaysia
Number of pages scraped: 20
Number of posts scraped: 1938
Link to dataset on Huggingface
cleaned_scraped_headlinesscraped_xsum1sfia-9-scraped
SFIA-9-Scraped Dataset
This repository contains the SFIA-9-Scraped dataset, a JSON collection of the Skills Framework for the Information Age (SFIA) version 9 categories and levels, scraped for non-commercial research use.
🚀 Dataset Overview
Name: SFIA-9-Scraped
Hugging Face: Programmer-RD-AI/sfia-9-scraped
DOI: 10.57967/hf/5746
Author: Ranuga Disansa Gamage
Revision: 89feeb8
Publisher: Hugging Face
Year: 2025
Use this dataset to build RAG systems, taxonomy-driven… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sfia-9-scraped.goodreads-scraped-datasetcypherhackz-scraped
website: cypherhackz
Number of pages scraped: 9
Number of posts scraped: 805
Link to dataset on Huggingface
scraped-forum-threadshealifyLLM-QA-scraped-dataset
