cleaned-data
HRM-Text-data-io-cleaned-20260515Pre-built HRM-Text pretraining dataset from raw data using the data_io cleaning scripts.
Citation
If you find this project or our paper useful, please consider citing our paper:
@misc{wang2026hrmtextefficientpretrainingscaling,
title={HRM-Text: Efficient Pretraining Beyond Scaling},
author={Guan Wang and Changling Liu and Chenyu Wang and Cai Zhou and Yuhao Sun and Yifei Wu and Shuai Zhen and Luca Scimeca and Yasin Abbasi Yadkori},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/sapientinc/HRM-Text-data-io-cleaned-20260515.booksummaries_cleanedGutenberg-BookCorpus-Cleaned-Data-English
Gutenberg-BookCorpus-Cleaned-Data-English
This dataset is been cleaned and preprocessed using Gutenberg_English_Preprocessor class method (given below) from preference Kaggle dataset 75,000+ Gutenberg Books and Metadata 2025. This dataset is only specialisation for english contented with rights as "Public domain in the USA" hence you can free used it anywhere.
Following reference metadata of Gutenberg is also available and downloaded it using following CLI command below :-
pip… See the full description on the dataset page: https://huggingface.co/datasets/incredible45/Gutenberg-BookCorpus-Cleaned-Data-English.cleaned_manga_dataset
Licensing and source material
This repository contains derived/processed images from publicly accessible manga
sources (GANMA! and Comic Walker / カドコミ). The original works are copyrighted by
their respective rights holders. No ownership of the original artwork is claimed.
The CC BY-NC 4.0 license applies only to the annotations and dataset metadata
created by the dataset authors, not to the underlying original artwork.
Unpacking the dataset
The dataset ships as a… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_manga_dataset.reddit_pushshift_dataset_cleaned
🚀 Cleaned Reddit Pushshift Dataset (Parquet)
📖 Dataset Description
This dataset is a highly optimized, meticulously cleaned, and structured version of the raw Reddit Pushshift dump. It has been transformed from massive .zst compressed JSON files into highly efficient, columnar Parquet files, making it immediately ready for Big Data analytics, SQL querying (via DuckDB), and Large Language Model (LLM) training.
The dataset includes both Reddit Submissions (Posts)… See the full description on the dataset page: https://huggingface.co/datasets/anhchanghoangsg/reddit_pushshift_dataset_cleaned.cleaned_webtoon_dataset
Licensing and source material
This repository contains derived/processed images from publicly accessible web comics/manga sources. The original works are copyrighted by their respective rights holders. No ownership of the original artwork is claimed.
The CC BY-NC 4.0 license applies only to the annotations and dataset metadata created by the dataset authors, not to the underlying original artwork.
Unpacking the dataset
The dataset ships as a zstd-compressed… See the full description on the dataset page: https://huggingface.co/datasets/Vasyanator2/cleaned_webtoon_dataset.
