datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blogspot_raw
Dataset Card for blogspot raw dataset
Dataset Summary
This dataset is a corpus of raw blogposts from blogspot mostly in the English language. It was obtained by scraping corpora of webarchive and commoncrawl.
Supported Tasks and Leaderboards
The dataset may be used for training language models or serve other research interests.
Languages
Mostly English language, but some outliers may occur.
Dataset Structure
Distribution
The distribution… See the full description on the dataset page: https://huggingface.co/datasets/mschi/blogspot_raw.portuguese-blogs
Dataset Details
Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text.
Texts are separated by the string <|endoftext|>.
Uses
Training language models.
Dataset Structure
A simple text file with articles separated by <|endoftext|> between each text.
Dataset Creation
First semester of 2024.
Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.ptbr-blogs
PT-BR Blogs (long-form, C4-derived)
Part of the MagTina350m pretrain corpus release by Dataseek
under the Magestic.ai brand. This is one of nine silver-layer datasets that fed
dataseek/magtina350m-base.
Summary
185 K long-form Brazilian-Portuguese blog posts (≥ 5 K words each) extracted from C4 by filtering Blogspot, WordPress, Medium and similar platform domains. Higher per-document quality than generic web; useful for stylistic diversity and long-context training.… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-blogs.azerbaijani-blogs
Azerbaijani Blogs dataset
Dataset Details
Dataset Description
This dataset provides blogs written in azerbaijani language with categories and tags for each.
Language(s) (NLP): Azerbaijani
License: Apache license 2.0
Data Source
All the data was found in public resources of kayzen.az blogging website without any restriction.
blogsetbr
BlogSet-BR
Reprodução do dataset BlogSet-BR criado pela universidade PUCRS.
Dataset Original
O dataset original (sem modificações) encontra-se em blogsetbr-original.csv (7.477.853
registros).
Dataset Modificado
Uma cópia modificada do dataset pode ser encontrada em blogsetbr-modificado.csv (7.468.541
registros). Foi modificado:
Remoção de registros duplicados e com problemas de escape (9.312 registros removidos).
Adicionado um cabeçalho ao arquivo.
O seguinte… See the full description on the dataset page: https://huggingface.co/datasets/tallesl/blogsetbr.llm-rag-agent-blogs
llm-rag-agent-blogs
Technical blogs on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline
Dataset Structure
This dataset contains three subsets:
llm: Large Language Model related content
rag: Retrieval-Augmented Generation related content
agent: AI Agent related content
Usage
from datasets import load_dataset
# Load all subsets
dataset = load_dataset("GXMZU/llm-rag-agent-blogs")
# Load specific subset
llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-blogs.netlab-blogs
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/kenshinx/netlab-blogs.
