blogs
Datasets
All datasets matching “blogs”AWS-White-Papers-and-Blogsblogsblogspot_raw
Dataset Card for blogspot raw dataset
Dataset Summary
This dataset is a corpus of raw blogposts from blogspot mostly in the English language. It was obtained by scraping corpora of webarchive and commoncrawl.
Supported Tasks and Leaderboards
The dataset may be used for training language models or serve other research interests.
Languages
Mostly English language, but some outliers may occur.
Dataset Structure
Distribution
The distribution… See the full description on the dataset page: https://huggingface.co/datasets/mschi/blogspot_raw.reddit-blogspot-twitterblogset-brportuguese-blogs
Dataset Details
Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text.
Texts are separated by the string <|endoftext|>.
Uses
Training language models.
Dataset Structure
A simple text file with articles separated by <|endoftext|> between each text.
Dataset Creation
First semester of 2024.
Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.
