CoolFace
20 results

Scrape

SKT-NRS /GIT-SCRAPED 🚀 SKT-NRS / GIT-SCRAPED This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets. Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures. 📂 Repository Structure All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.1K<n<10K1 likes4.4k downloads3mo agoHugging Facehudsongouge /AoPS-Scrape AoPS-Scrape Problems and solutions scraped from Art of Problem Solving (AoPS) Online class homework endpoints. Obtained legally in accordance with AoPS's Terms of Service. This is not unauthorized redistribution of pirated material — access was through a legitimate authenticated AoPS Online class session. Splits Splits are named by scrape date (YYYY_MM_DD), plus a cross-date content-deduplicated split: Split Rows Notes deduplicated 29,964 One row per… See the full description on the dataset page: https://huggingface.co/datasets/hudsongouge/AoPS-Scrape.tabularquestion-answering10K<n<100K2 likes4.2k downloads2mo agoHugging FaceGugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.8k downloads18d agoHugging Facethuerey-group /apebench-scraped APEBench Scraped A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0. ⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.1 likes1.2k downloads2y agoHugging FaceAbstractPhil /IMDB-PUBLIC-SCRAPED Hello World with Hugging Face Current Date: 2025-03-19 04:36:42.698271 So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later. It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks. I'll be working out the problems and getting the scraper working correctly at some point soon. 1 likes908 downloads4mo agoHugging Facesayurio /english-manhwa-scrape English Manhwa Scrape Dataset Request More ScrapesOrder Private Scrapes Note: The total files are about 150+GB and my internet connection is slow af. So I'll be uploading them in batches File Structure: files/ - Manhwa 1 Name - Chapter 1.cbz - Chapter 2.cbz - Chapter 3.cbz - ...... - Chapter n.cbz - Manhwa 2 Name - Chapter 1.cbz - Chapter 2.cbz - Chapter 3.cbz - ...... - Chapter n.cbz Due to maximum 10000 files in a repo limit of huggingface, I had to further… See the full description on the dataset page: https://huggingface.co/datasets/sayurio/english-manhwa-scrape.image-to-text100K<n<1M2 likes789 downloads6mo agoHugging Face