datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.paper2env-scraped
Paper2Env — scraped arXiv
One row per task. Each task is a paper-reproduction subtask with a verification
script (verify.sh) that scores submissions, plus a text-only git diff
patch against an upstream GitHub repo at a pinned commit.
Per-task binary artefacts (paper PDF, assets, expected outputs for grading,
binary file additions to the student repo) live in the companion repo
thibble/paper2env-artifacts under
scraped/<paper_id>/<task_id>.tar.gz.
Reconstruct a task… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-scraped.
