datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.scraped-chatgpt-conversations
Dataset Card for Dataset Name
Dataset Summary
scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.paper2env-scraped
Paper2Env — scraped arXiv
One row per task. Each task is a paper-reproduction subtask with a verification
script (verify.sh) that scores submissions, plus a text-only git diff
patch against an upstream GitHub repo at a pinned commit.
Per-task binary artefacts (paper PDF, assets, expected outputs for grading,
binary file additions to the student repo) live in the companion repo
thibble/paper2env-artifacts under
scraped/<paper_id>/<task_id>.tar.gz.
Reconstruct a task… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-scraped.
