CoolFace
17 results

cleaned-data

sapientinc /HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts. Citation If you find this project or our paper useful, please consider citing our paper: @misc{wang2026hrmtextefficientpretrainingscaling, title={HRM-Text: Efficient Pretraining Beyond Scaling}, author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.texttext-generation100M<n<1B17 likes3.7k downloads4mo agoHugging Faceada-datadruids /booksummaries_cleanedtext10K<n<100K0 likes1.9k downloads2y agoHugging Faceincredible45 /Gutenberg-BookCorpus-Cleaned-Data-English Gutenberg-BookCorpus-Cleaned-Data-English This dataset is been cleaned and preprocessed using Gutenberg_English_Preprocessor class method (given below) from preference Kaggle dataset 75,000+ Gutenberg Books and Metadata 2025. This dataset is only specialisation for english contented with rights as "Public domain in the USA" hence you can free used it anywhere. Following reference metadata of Gutenberg is also available and downloaded it using following CLI command below :- pip… See the full description on the dataset page: https://huggingface.co/datasets/incredible45/Gutenberg-BookCorpus-Cleaned-Data-English.text10K<n<100K15 likes1.3k downloads1y agoHugging FaceVasyanator2 /cleaned_manga_dataset Licensing and source material This repository contains derived/processed images from publicly accessible manga sources (GANMA! and Comic Walker / カドコミ). The original works are copyrighted by their respective rights holders. No ownership of the original artwork is claimed. The CC BY-NC 4.0 license applies only to the annotations and dataset metadata created by the dataset authors, not to the underlying original artwork. Unpacking the dataset The dataset ships as a… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_manga_dataset.imageimage-to-imagen<1K0 likes768 downloads3mo agoHugging Faceanhchanghoangsg /reddit_pushshift_dataset_cleaned 🚀 Cleaned Reddit Pushshift Dataset (Parquet) 📖 Dataset Description This dataset is a highly optimized, meticulously cleaned, and structured version of the raw Reddit Pushshift dump. It has been transformed from massive .zst compressed JSON files into highly efficient, columnar Parquet files, making it immediately ready for Big Data analytics, SQL querying (via DuckDB), and Large Language Model (LLM) training. The dataset includes both Reddit Submissions (Posts)… See the full description on the dataset page: https://huggingface.co/datasets/anhchanghoangsg/reddit_pushshift_dataset_cleaned.3 likes580 downloads6mo agoHugging FaceVasyanator2 /cleaned_webtoon_dataset Licensing and source material This repository contains derived/processed images from publicly accessible web comics/manga sources. The original works are copyrighted by their respective rights holders. No ownership of the original artwork is claimed. The CC BY-NC 4.0 license applies only to the annotations and dataset metadata created by the dataset authors, not to the underlying original artwork. Unpacking the dataset The dataset ships as a zstd-compressed… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_webtoon_dataset.imageimage-to-imagen<1K0 likes483 downloads3mo agoHugging Face