CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fabiovilao /portuguese-blogs Dataset Details Blog-1 may include other languages in an unstructured text format without markdown. The latest one, Blog-6, is formatted in markdown and may contain less other languages text. Texts are separated by the string <|endoftext|>. Uses Training language models. Dataset Structure A simple text file with articles separated by <|endoftext|> between each text. Dataset Creation First semester of 2024. Bias, Risks, and Limitations… See the full description on the dataset page: https://huggingface.co/datasets/fabiovilao/portuguese-blogs.texttext-generation100M<n<1B0 likes82 downloads2y agoHugging Face02dataseek /ptbr-blogs PT-BR Blogs (long-form, C4-derived) Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 185 K long-form Brazilian-Portuguese blog posts (≥ 5 K words each) extracted from C4 by filtering Blogspot, WordPress, Medium and similar platform domains. Higher per-document quality than generic web; useful for stylistic diversity and long-context training.… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-blogs.tabulartext-generation100K<n<1M1 likes66 downloads5mo agoHugging Face03aznlp /azerbaijani-blogs Azerbaijani Blogs dataset Dataset Details Dataset Description This dataset provides blogs written in azerbaijani language with categories and tags for each. Language(s) (NLP): Azerbaijani License: Apache license 2.0 Data Source All the data was found in public resources of kayzen.az blogging website without any restriction. texttext-classification1K<n<10K3 likes38 downloads2y agoHugging Face04GXMZU /llm-rag-agent-blogs llm-rag-agent-blogs Technical blogs on LLM, RAG, and AI Agents - Knowledge base for RAG pipeline Dataset Structure This dataset contains three subsets: llm: Large Language Model related content rag: Retrieval-Augmented Generation related content agent: AI Agent related content Usage from datasets import load_dataset # Load all subsets dataset = load_dataset("GXMZU/llm-rag-agent-blogs") # Load specific subset llm_data =… See the full description on the dataset page: https://huggingface.co/datasets/GXMZU/llm-rag-agent-blogs.texttext-generation1K<n<10K1 likes22 downloads9mo agoHugging Face05kenshinx /netlab-blogs Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/kenshinx/netlab-blogs.texttext-generationn<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.