CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceTB /openstax_paragraphsTexbooks from openstax.org with their chapters, abstracts and sections. Sample: { "book_title":"World History Volume 1, to 1500", "language":"en", "chapters":[ { "title":"Preface", "abstract":"None", "sections":[ { "title":"About OpenStax", "paragraph":"OpenStax is part of Rice University, which is a 501(c)(3) nonprofit..." }, { "title":"About OpenStax Resources"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/openstax_paragraphs.textn<1K6 likes2.9k downloads3y agoHugging Face02agentlans /wikipedia-paragraphs Wikipedia Paragraph Samples Dataset Description This dataset contains paragraphs extracted from randomly selected English Wikipedia articles. It provides a diverse sample of Wikipedia content across various topics. Dataset Details Name: Wikipedia Paragraph Samples Version: 1.0 Date Created: 2024-08-20 Language: English Format: JSONLines Contents Each line in the dataset represents a single paragraph and contains two fields: Title of the Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs.texttext-classification10K<n<100K3 likes419 downloads2y agoHugging Face03pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes136 downloads2y agoHugging Face04pere /wiki_paragraphs_english WIKI Paragraphs English A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.tabulartext-generation1M<n<10M0 likes87 downloads2y agoHugging Face05agentlans /wikipedia-paragraphs-complete Wikipedia Paragraphs Complete Dataset This dataset consists of English Wikipedia paragraphs ranging from 1 000 to 8 000 characters in length. It was sourced from the Wikimedia dump: "wikimedia/wikipedia", "20231101.en". Preprocessing Steps The dataset has undergone extensive cleaning and normalization, including: Removing brackets Removing HTML tags Normalizing bullet points, hyphenated words, quotation marks, Unicode characters, and whitespace Replacing email… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-complete.texttext-generation1M<n<10M1 likes72 downloads1y agoHugging Face06yimingwang123 /grade_labeled_wiki_paragraphs Grade-Labeled Wiki Paragraphs (GPT-4.1 Nano) This dataset contains Wikipedia paragraphs simplified to different grade reading levels (targeting Grade 1-12) using the GPT-4.1 Nano model. Dataset Description Dataset Summary The dataset consists of pairs of original Wikipedia paragraphs and their machine-generated simplified versions. The simplification aims to make the text understandable for readers at specific US grade levels while preserving the core… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade_labeled_wiki_paragraphs.tabular10K<n<100K0 likes27 downloads1y agoHugging Face07agentlans /wikipedia-paragraphs-direct-paraphrases Wikipedia Paragraphs Direct Paraphrases Paraphrases of Wikipedia paragraphs using AI large language models. Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt: Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.texttext-generation10K<n<100K0 likes20 downloads3mo agoHugging Face08AL49 /lotr_paragraphstext1K<n<10K1 likes14 downloads3y agoHugging Face09vishwap1991 /openstax_paragraphsTexbooks from openstax.org with their chapters, abstracts and sections. Sample: { "book_title":"World History Volume 1, to 1500", "language":"en", "chapters":[ { "title":"Preface", "abstract":"None", "sections":[ { "title":"About OpenStax", "paragraph":"OpenStax is part of Rice University, which is a 501(c)(3) nonprofit..." }, { "title":"About OpenStax Resources"… See the full description on the dataset page: https://huggingface.co/datasets/vishwap1991/openstax_paragraphs.textn<1K0 likes8 downloads5mo agoHugging Face10Mathieu-Thomas-JOSSET /michael_paragraphs_units_5 michael_paragraphs_units_5 Dataset uploaded with Python via huggingface_hub. Files JSONL source file uploaded to this repository Notes Custom dataset Uploaded automatically from a local file tabulartext-generation10K<n<100K0 likes6 downloads7mo agoHugging Face11Rileyw02 /Synthetic_Clinical_Notes_Paragraphstext10K<n<100K0 likes4 downloads8mo agoHugging Face12genaforvena /huivam_finnegans_wake_paragraphs Dataset Overview Dataset Name: huivam_finnegans_wake_paragraphs Creator: Platform: Hugging Face Datasets Dataset Context Source Material: Likely derived from "Finnegans Wake", a novel by James Joyce known for its experimental language and complex structure. Content Type: Paragraphs (text format) Expected Dataset Details (Not Selected, Assumed from Typical Dataset Pages) Key Features: Text Data: Paragraphs from Finnegans Wake Language: English… See the full description on the dataset page: https://huggingface.co/datasets/genaforvena/huivam_finnegans_wake_paragraphs.textn<1K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.