CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SKT-NRS /GIT-SCRAPED 🚀 SKT-NRS / GIT-SCRAPED This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets. Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures. 📂 Repository Structure All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.1K<n<10K1 likes4.4k downloads3mo agoHugging Face02Gugu8 /Scraped-Datagated Scraped-Data — Building to 1 TB A continuously-growing, cleaned text corpus scraped, parsed, cleaned, losslessly-zipped, and auto-pushed from a CPU-only Google Colab pipeline. Current repo size: ~26 GB → target: 1 TB (incremental batch uploads). Current files File Size Contents corpus_batch1.clean.zip 4.80 GB ~2.9M Wikipedia docs (stored, from first scrape) corpus_batch2.clean.zip 4.85 GB ~2.5M docs corpus_batch3.clean.zip 4.91 GB ~2.3M docs… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Scraped-Data.texttext-generation100M<n<1B0 likes1.7k downloads21d agoHugging Face03thuerey-group /apebench-scraped APEBench Scraped A representative subset of datasets created using the APEBench benchmark suite using version 0.1.0. ⚠️ Note that APEBench is designed to procedurally generate all its training and test data. This allows for advanced features like benchmarking approaches with differentiable physics. Hence, there is no need to download this dataset as it can be easily re-generated using APEBench which can be installed via pip install apebench. See also here for how to scrape datasets.… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped.1 likes1.2k downloads2y agoHugging Face04AbstractPhil /IMDB-PUBLIC-SCRAPED Hello World with Hugging Face Current Date: 2025-03-19 04:36:42.698271 So this one didn't quite finish scraping. I'll fix the software and rerun the scraping later. It had some flaws with the multithreading where it would upload the same archives and overwrite the originals, which caused annoying problems and quirks. I'll be working out the problems and getting the scraper working correctly at some point soon. 1 likes899 downloads4mo agoHugging Face05thuerey-group /apebench-scraped-old APEBench-scraped (old) All datasets scraped from the APEBench benchmark suite with the version used for the Neurips submission. Download Download without large files GIT_LFS_SKIP_SMUDGE=1 git clone git@hf.co:datasets/thuerey-group/apebench-scraped-old Afterwards, you can inspect the repository and download the files you need. For example, for 1d_diff_adv: git lfs install git lfs pull -I "data/1d_diff_adv*" Alternatively, you can download the entire repository with large… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/apebench-scraped-old.0 likes769 downloads2y agoHugging Face06Asib27 /github_repo_scrapedScraped python packages from github. 0 likes703 downloads2y agoHugging Face07mondk /vi.wikipedia-scraped-dataSorry for not having English. bru 4 likes307 downloads2mo agoHugging Face08Isamu136 /penetration_testing_scraped_dataset Dataset Card for "penetration_testing_scraped_dataset" More Information needed text100K<n<1M14 likes257 downloads3y agoHugging Face09Sinfsr /telegram-scraped-data0 likes190 downloads7mo agoHugging Face10barryallen16 /fitcheck-scraped-v1image100K<n<1M0 likes180 downloads8mo agoHugging Face11ar852 /scraped-chatgpt-conversations Dataset Card for Dataset Name Dataset Summary scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.question-answering100K<n<1M12 likes115 downloads3y agoHugging Face12Nicolas-BZRD /English_French_Webpages_Scraped_Translated English French Webpages Scraped Translated Dataset Summary French/English parallel texts for training translation models. Over 17.1 million sentences in French and English. Dataset created by Chris Callison-Burch, who crawled millions of web pages and then used a set of simple heuristics to transform French URLs onto English URLs, and assumed that these documents are translations of each other. This is the main dataset of Workshop on Statistical Machine Translation (WML)… See the full description on the dataset page: https://huggingface.co/datasets/Nicolas-BZRD/English_French_Webpages_Scraped_Translated.texttranslation10M<n<100M3 likes98 downloads3y agoHugging Face13mondk /arena.ai-code-leaderboard-scrapedlink: https://arena.ai/leaderboard/code/ tabularn<1K0 likes98 downloads29d agoHugging Face14DarkforumsHeiter /Darkforums-Scraped-Dump-osint-Datatextn<1K1 likes88 downloads3mo agoHugging Face15IrishHotelReviews /scraped-hotel-reviewstabular10K<n<100K0 likes86 downloads4mo agoHugging Face16FinancialSupport /ScrapedJobstabularn<1K0 likes64 downloads3y agoHugging Face17narteybrown /sobolo-corpus-multilang-scraped0 likes60 downloads7mo agoHugging Face18OpenPipe /hacker-news-scraped-storiestabular1M<n<10M1 likes55 downloads2y agoHugging Face19AIMO-Corpus /PolyMath-Scraped-Raw PolyMath Scraped PolyMath is a curated dataset of 11,090 high-difficulty mathematical problems designed for training reasoning models. Built for the AIMO Math Corpus Prize. Existing math datasets (NuminaMath-1.5, OpenMathReasoning) suffer from high noise rates in their hardest samples and largely unusable proof-based problems. PolyMath addresses both issues through: Data scraping: problems sourced from official competition PDFs absent from popular datasets, using a… See the full description on the dataset page: https://huggingface.co/datasets/AIMO-Corpus/PolyMath-Scraped-Raw.text10K<n<100K0 likes55 downloads8mo agoHugging Face20LuxWorld /scraped-episodestext1K<n<10K0 likes51 downloads2y agoHugging Face21OpenTransformer /scraped-web-datatext100K<n<1M0 likes51 downloads9mo agoHugging Face22asingh15 /NYT-Connections-Verl-Scrapedtext1K<n<10K0 likes49 downloads5mo agoHugging Face23haizad /jurnal-malaysia-scraped website: jurnal-malaysia Number of pages scraped: 20 Number of posts scraped: 1938 Link to dataset on Huggingface text1K<n<10K0 likes47 downloads3y agoHugging Face24Jc710 /cleaned_scraped_headlinestext-classificationn<1K0 likes46 downloads9mo agoHugging Face25mpalaval /scraped_xsum1textn<1K0 likes45 downloads3y agoHugging Face26Programmer-RD-AI /sfia-9-scraped SFIA-9-Scraped Dataset This repository contains the SFIA-9-Scraped dataset, a JSON collection of the Skills Framework for the Information Age (SFIA) version 9 categories and levels, scraped for non-commercial research use. 🚀 Dataset Overview Name: SFIA-9-Scraped Hugging Face: Programmer-RD-AI/sfia-9-scraped DOI: 10.57967/hf/5746 Author: Ranuga Disansa Gamage Revision: 89feeb8 Publisher: Hugging Face Year: 2025 Use this dataset to build RAG systems, taxonomy-driven… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sfia-9-scraped.textn<1K2 likes45 downloads1y agoHugging Face27Johnvictor1275 /goodreads-scraped-dataset0 likes43 downloads6mo agoHugging Face28haizad /cypherhackz-scraped website: cypherhackz Number of pages scraped: 9 Number of posts scraped: 805 Link to dataset on Huggingface textn<1K0 likes42 downloads3y agoHugging Face29dolores-haze /scraped-forum-threadstextn<1K0 likes42 downloads1y agoHugging Face30falan42 /healifyLLM-QA-scraped-datasettext1K<n<10K1 likes41 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.