scraped
2023_Scraped_SDXLSDXL_Scraped_Loras_2024exp_24_sft-julia_sft_reverse_instruct_for_scraped_funcssft_16bit_vllmscraped_data_v13_urls_Llama-2-7b-lr-2e-5-ep3scraped_data_v13_urls_gpt2-xl-lr-2e-5-ep3scraped_data_v13_urls_Phi-3-mini-4k-instruct-lr-5e6-ep3scraped_data_v13_urls_gpt2-lr-2e-5-ep3scraped_data_v13_urls_Phi-3-mini-4k-instruct-ep3
Datasets
All datasets matching “scraped”GIT-SCRAPED
🚀 SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
📂 Repository Structure
All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.Scraped-Data
Scraped-Data — Building to 1 TB
A continuously-growing, cleaned text corpus scraped, parsed, cleaned,
losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline.
Current repo size: ~26 GB → target: 1 TB (incremental batch uploads).
Current files
File
Size
Contents
corpus_batch1.clean.zip
4.80 GB
~2.9M Wikipedia docs (stored, from first scrape)
corpus_batch2.clean.zip
4.85 GB
~2.5M docs
corpus_batch3.clean.zip
4.91 GB
~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.apebench-scraped
APEBench Scraped
A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0.
⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.IMDB-PUBLIC-SCRAPED
Hello World with Hugging Face
Current Date: 2025-03-19 04:36:42.698271
So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later.
It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks.
I'll be working out the problems and getting the scraper working correctly at some point soon.
apebench-scraped-old
APEBench-scraped (old)
All datasets scraped from the APEBench benchmark suite with the version used
for the Neurips
submission.
Download
Download without large files
GIT_LFS_SKIP_SMUDGE=1 git clone git@hf.co:datasets/thuerey-group/apebench-scraped-old
Afterwards, you can inspect the repository and download the files you need. For
example, for 1d_diff_adv:
git lfs install
git lfs pull -I "data/1d_diff_adv*"
Alternatively, you can download the entire repository with large… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped-old.github_repo_scrapedScraped python packages from github.
