CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mschi /blogspot_raw Dataset Card for blogspot raw dataset Dataset Summary This dataset is a corpus of raw blogposts from blogspot mostly in the English language. It was obtained by scraping corpora of webarchive and commoncrawl. Supported Tasks and Leaderboards The dataset may be used for training language models or serve other research interests. Languages Mostly English language, but some outliers may occur. Dataset Structure Distribution The distribution… See the full description on the dataset page: https://huggingface.co/datasets/mschi/blogspot_raw.imagetext-classificationn<1K1 likes105 downloads4y agoHugging Face02fabiovilao /portuguese-blogs Dataset Details Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text. Texts are separated by the string <|endoftext|>. Uses Training language models. Dataset Structure A simple text file with articles separated by <|endoftext|> between each text. Dataset Creation First semester of 2024. Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.texttext-generation100M<n<1B0 likes82 downloads2y agoHugging Face03dataseek /ptbr-blogs PT-BR Blogs (long-form, C4-derived) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 185 K long-form Brazilian-Portuguese blog posts (≥ 5 K words each) extracted from C4 by filtering Blogspot, WordPress, Medium and similar platform domains. Higher per-document quality than generic web; useful for stylistic diversity and long-context training.… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-blogs.tabulartext-generation100K<n<1M1 likes66 downloads5mo agoHugging Face04aznlp /azerbaijani-blogs Azerbaijani Blogs dataset Dataset Details Dataset Description This dataset provides blogs written in azerbaijani language with categories and tags for each. Language(s) (NLP): Azerbaijani License: Apache license 2.0 Data Source All the data was found in public resources of kayzen.az blogging website without any restriction. texttext-classification1K<n<10K3 likes38 downloads2y agoHugging Face05tallesl /blogsetbr BlogSet-BR Reprodução do dataset BlogSet-BR criado pela universidade PUCRS. Dataset Original O dataset original (sem modificações) encontra-se em blogsetbr-original.csv (7.477.853 registros). Dataset Modificado Uma cópia modificada do dataset pode ser encontrada em blogsetbr-modificado.csv (7.468.541 registros). Foi modificado: Remoção de registros duplicados e com problemas de escape (9.312 registros removidos). Adicionado um cabeçalho ao arquivo. O seguinte… See the full description on the dataset page: https://huggingface.co/datasets/tallesl/blogsetbr.text-generation1M<n<10M0 likes23 downloads2y agoHugging Face06GXMZU /llm-rag-agent-blogs llm-rag-agent-blogs Technical blogs on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline Dataset Structure This dataset contains three subsets: llm: Large Language Model related content rag: Retrieval-Augmented Generation related content agent: AI Agent related content Usage from datasets import load_dataset # Load all subsets dataset = load_dataset("GXMZU/llm-rag-agent-blogs") # Load specific subset llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-blogs.texttext-generation1K<n<10K1 likes22 downloads9mo agoHugging Face07kenshinx /netlab-blogs Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/kenshinx/netlab-blogs.texttext-generationn<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.