CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agentlans /wikipedia-paragraphs Wikipedia Paragraph Samples Dataset Description This dataset contains paragraphs extracted from randomly selected English Wikipedia articles. It provides a diverse sample of Wikipedia content across various topics. Dataset Details Name: Wikipedia Paragraph Samples Version: 1.0 Date Created: 2024-08-20 Language: English Format: JSONLines Contents Each line in the dataset represents a single paragraph and contains two fields: Title of the Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs.texttext-classification10K<n<100K3 likes419 downloads2y agoHugging Face02agentlans /wikipedia-first-paragraphtexttext-classification10M<n<100M0 likes175 downloads1y agoHugging Face03pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes136 downloads2y agoHugging Face04pere /wiki_paragraphs_english WIKI Paragraphs English A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.tabulartext-generation1M<n<10M0 likes87 downloads2y agoHugging Face05lumasik /Synthetic-Pretrain-Paragraphs-150Topics Synthetic-Pretrain-Paragraphs-150Topics A synthetic dataset consisting of continuous text paragraphs in Russian and English, generated using Qwen2.5-7B-Instruct. Dataset Curation Source: Generated via vLLM on an RTX 3060. Topics: Covers 150 fundamental fields of knowledge, ranging from programming to speleology. Constraint: Strict prohibition on greetings, lists, and conversational fillers; contains only raw factual descriptions. Warning: As pure synthetic data, it… See the full description on the dataset page: https://huggingface.co/datasets/lumasik/Synthetic-Pretrain-Paragraphs-150Topics.texttext-generation100K<n<1M2 likes77 downloads5mo agoHugging Face06agentlans /wikipedia-paragraphs-complete Wikipedia Paragraphs Complete Dataset This dataset consists of English Wikipedia paragraphs ranging from 1 000 to 8 000 characters in length. It was sourced from the Wikimedia dump: "wikimedia/wikipedia", "20231101.en". Preprocessing Steps The dataset has undergone extensive cleaning and normalization, including: Removing brackets Removing HTML tags Normalizing bullet points, hyphenated words, quotation marks, Unicode characters, and whitespace Replacing email… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-complete.texttext-generation1M<n<10M1 likes72 downloads1y agoHugging Face07agentlans /wikipedia-paragraph-summaries Wikipedia Paragraph Summaries Dataset The Wikipedia Paragraph Summaries Dataset is designed for the task of text summarization, specifically generating concise summaries from paragraphs extracted from English Wikipedia articles. Each entry in the dataset consists of an input paragraph and its corresponding summary, facilitating research in natural language processing (NLP) and machine learning. Data Format: The dataset is provided in JSON Lines format (.jsonl), where each line… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-summaries.textsummarization10K<n<100K0 likes39 downloads2y agoHugging Face08schneiderkamplab /dfm12-reordering-integrated-nl-paragraph-reordering dfm12-reordering-integrated-nl-paragraph-reordering Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-nl-paragraph-reordering.texttext-generation10K<n<100K0 likes30 downloads1d agoHugging Face09schneiderkamplab /dfm12-reordering-integrated-nn-paragraph-reordering dfm12-reordering-integrated-nn-paragraph-reordering Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-nn-paragraph-reordering.texttext-generation10K<n<100K0 likes28 downloads1d agoHugging Face10schneiderkamplab /dfm12-reordering-integrated-nb-paragraph-reordering dfm12-reordering-integrated-nb-paragraph-reordering Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-nb-paragraph-reordering.texttext-generation10K<n<100K0 likes26 downloads1d agoHugging Face11agentlans /wikipedia-paragraph-sft Wikipedia Paragraph Supervised Finetuning Dataset Model Description This dataset is designed for training language models to generate supervised finetuning data from raw text. It consists of text passages and corresponding question-answer pairs in JSONLines format. Intended Use The primary purpose of this dataset is to enable large language models (LLMs) to generate high-quality supervised finetuning data from raw text inputs, useful for creating custom… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-sft.textquestion-answering10K<n<100K1 likes25 downloads2y agoHugging Face12schneiderkamplab /dfm12-reordering-integrated-sv-paragraph-reordering dfm12-reordering-integrated-sv-paragraph-reordering Published accepted-only DFM12 subset. Local audit-snapshot fields describe the pre-publication build, not Hub publication status. Only completed kept decisions with all three scores at least 4 are included, after deterministic gates. Automated review is not native-speaker certification. Exclusion metadata contains only IDs/status/errors/scores/reasons, never excluded conversations. Full native messages and explicit assistant… See the full description on the dataset page: https://huggingface.co/datasets/schneiderkamplab/dfm12-reordering-integrated-sv-paragraph-reordering.texttext-generation10K<n<100K0 likes22 downloads1d agoHugging Face13agentlans /wikipedia-paragraphs-direct-paraphrases Wikipedia Paragraphs Direct Paraphrases Paraphrases of Wikipedia paragraphs using AI large language models. Paragraphs from agentlans/wikipedia-paragraphs-complete sample_k10000 and sample_k50000 splits Paraphrased using Qwen/Qwen3.5-9B and a distilled Qwen/Qwen3-4B-Instruct-2507 with the following prompt: Rewrite the following paragraph entirely in your own words while preserving every fact, detail, meaning, nuance, and level of specificity. Do not add, remove, reinterpret… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraphs-direct-paraphrases.texttext-generation10K<n<100K0 likes18 downloads3mo agoHugging Face14Lots-of-LoRAs /task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph.texttext-generationn<1K0 likes16 downloads2y agoHugging Face15Shakil2448868 /bangla-word-to-paragraphtexttext-generation1K<n<10K0 likes10 downloads2y agoHugging Face16YawKar /wikitext_with_entitled_paragraphstexttext-generation100K<n<1M0 likes9 downloads3y agoHugging Face17agentlans /wikipedia-paragraph-conversation Wikipedia Paragraphs Conversation Dataset This dataset contains approximately 800 high-quality English Wikipedia paragraphs paired with synthetic, multi-turn conversations. It is specifically designed to train or fine-tune LLMs to act as synthetic data generators—transforming static knowledge into natural, multi-turn dialogue. Dataset Summary The goal of this dataset is to bridge the gap between raw factual prose and interactive conversational formats. Each row features a… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/wikipedia-paragraph-conversation.texttext-generation1K<n<10K0 likes8 downloads4mo agoHugging Face18Mathieu-Thomas-JOSSET /michael_paragraphs_units_5 michael_paragraphs_units_5 Dataset uploaded with Python via huggingface_hub. Files JSONL source file uploaded to this repository Notes Custom dataset Uploaded automatically from a local file tabulartext-generation10K<n<100K0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.