CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01singletongue /wikipedia-paragraphs wikipedia-paragraphs wikipedia-paragraphs is a dataset generated from Wikipedia, designed for natural language processing (NLP) research. Each entry contains cleaned paragraph text and Wikilink information extracted from a Wikipedia page, along with useful metadata such as categories, templates, and the associated Wikidata QID. Dataset structure Configurations The dataset is organized into multiple configurations, such as enwiki-20260607-v1.2.1.… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/wikipedia-paragraphs.tabular100M<n<1B2 likes10k downloads3mo agoHugging Face02JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids Thoughts for next-token prediction on k=8 token chunks of JackHsieh/statML-arxiv-40M-20M, generated by Qwen/Qwen3-4B in thinking mode. The prompted task: reason about what comes IMMEDIATELY next — the next k=8 tokens after the cut — and answer with a single unconstrained paragraph of dense reasoning, focused on the exact state at the cut and what the local grammar, notation, or argument forces next. Both the… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.tabular10M<n<100M0 likes435 downloads19d agoHugging Face03JackHsieh /4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained 4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained A pre-tokenized, tag-wrapped variant of JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids. The source thoughts are two parts: a <think> block, then a single paragraph of dense reasoning about the immediate continuation. Only the part after </think> — the paragraph — becomes the VALUE. The reasoning inside the think block is dropped. The VALUE is capped at 512 tokens.… See the full description on the dataset page: https://huggingface.co/datasets/JackHsieh/4B-think-paragraph.stride-train1-test32.k-8.statml-arxiv.qwen3-ids.kv-tags-explained.tabular10M<n<100M0 likes365 downloads19d agoHugging Face04mlfoundations-dev /d1_code_long_paragraphstabular10K<n<100K0 likes246 downloads1y agoHugging Face05pere /wiki_paragraphs_norwegian WIKI Paragraphs Norwegian A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_norwegian.tabulartext-generation1M<n<10M0 likes136 downloads2y agoHugging Face06oshizo /japanese-wikipedia-paragraphsA slightly modified version of the parsing and chunking method for singletongue/wikipedia-utils. Pre-processing was performed using oshizo/wikipedia-utils, which is a fork of the original repository, singletongue/wikipedia-utils. The Wikipedia data was crawled between 2023/12/5 and 2023/12/8. tabular10M<n<100M4 likes128 downloads3y agoHugging Face07skeskinen /books3_lowgrade_paragraphs Dataset Card for "books3_lowgrade_paragraphs" the_pile books3, books with smog grade difficulty estimate between 6.6 or and 7.1. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc. For easier books, see books3_basic_paragraphs tabular10M<n<100M0 likes108 downloads3y agoHugging Face08ulab-ai /arxiv-paragraphstabular1M<n<10M0 likes108 downloads11mo agoHugging Face09LumberChunker /GutenQA_Paragraphs 📚 GutenQA-Paragraphs GutenQA-Paragraphs consists on the same 100 Public Domain Narrative Books used in GutenQA. In this version, passages are extracted at the paragraph level. The GutenQA dataset, as available, is the result of applying the text segmentation method LumberChunker to the GutenQA-Paragraphs. The dataset is organized into the following columns: Book Name: The title of the book from which the passage is extracted. Book ID: A unique integer identifier assigned to each… See the full description on the dataset page: https://huggingface.co/datasets/LumberChunker/GutenQA_Paragraphs.tabularquestion-answering100K<n<1M0 likes103 downloads2y agoHugging Face10pere /wiki_paragraphs_english WIKI Paragraphs English A multi-split dataset for machine learning research and evaluation, containing text samples in JSON Lines format. Features Multiple splits for different use cases Random shuffle with Fisher-Yates algorithm Structured format with text and metadata Size-varied validation/test sets (100 to 10k samples) Splits Overview Split Name Samples Typical Usage train 1,000,000 Primary training data validation 10,000 Standard validation… See the full description on the dataset page: https://huggingface.co/datasets/pere/wiki_paragraphs_english.tabulartext-generation1M<n<10M0 likes87 downloads2y agoHugging Face11llm-book /jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3 Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3" More Information needed tabular1M<n<10M1 likes86 downloads3y agoHugging Face12mlfoundations-dev /d1_math_long_paragraphstabular10K<n<100K0 likes81 downloads1y agoHugging Face13ulab-ai /ResearchArcade-arxiv-paragraphstabular1M<n<10M0 likes57 downloads11mo agoHugging Face14skeskinen /books3_basic_paragraphs Dataset Card for "books3_basic_paragraphs" the_pile books3, books with smog grade difficulty estimate of 6.5 or under. Split into paragraphs and filtered out most 'non-paragraphs' like titles, tables of content, etc. tabular1M<n<10M0 likes55 downloads3y agoHugging Face15raminass /paragraphs Dataset Card for "paragraphs" More Information needed tabular10K<n<100K0 likes39 downloads3y agoHugging Face16GuillermoTBB /gp-long-paragraphstabular10M<n<100M1 likes27 downloads2y agoHugging Face17yimingwang123 /grade_labeled_wiki_paragraphs Grade-Labeled Wiki Paragraphs (GPT-4.1 Nano) This dataset contains Wikipedia paragraphs simplified to different grade reading levels (targeting Grade 1-12) using the GPT-4.1 Nano model. Dataset Description Dataset Summary The dataset consists of pairs of original Wikipedia paragraphs and their machine-generated simplified versions. The simplification aims to make the text understandable for readers at specific US grade levels while preserving the core… See the full description on the dataset page: https://huggingface.co/datasets/yimingwang123/grade_labeled_wiki_paragraphs.tabular10K<n<100K0 likes27 downloads1y agoHugging Face18ulab-ai /arxiv-paragraph-tablestabular100K<n<1M0 likes23 downloads11mo agoHugging Face19d0rj /RuBQ_2.0-paragraphs RuBQ_2.0-paragraphs For test and dev data see d0rj/RuBQ_2.0 tabularquestion-answering10K<n<100K2 likes22 downloads3y agoHugging Face20llm-book /jawiki-paragraphs Dataset Card for llm-book/jawiki-paragraphs 書籍『大規模言語モデル入門』で使用する Wikipedia 段落のデータセットです。 GitHub リポジトリ singletongue/wikipedia-utils で公開されているデータセットを利用しています。 Licence 本データセットで使用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。 tabular1M<n<10M0 likes21 downloads3y agoHugging Face21mlfoundations-dev /d1_science_long_paragraphstabular10K<n<100K0 likes20 downloads1y agoHugging Face22mlfoundations-dev /d1_code_long_paragraphs_10ktabular1K<n<10K0 likes19 downloads1y agoHugging Face23ulab-ai /ResearchArcade-arxiv-paragraph-figurestabular1M<n<10M0 likes19 downloads11mo agoHugging Face24ulab-ai /arxiv-paragraph-figurestabular1M<n<10M0 likes17 downloads11mo agoHugging Face25ulab-ai /ResearchArcade-arxiv-paragraph-citationstabular1M<n<10M0 likes17 downloads11mo agoHugging Face26Brianverb /secrethackatondata_paragraphstabularn<1K0 likes15 downloads4y agoHugging Face27SinclairSchneider /Bundestagsreden_Paragraphs_with_likes Dataset Card for "Bundestagsreden_Paragraphs_with_likes" More Information needed tabular100K<n<1M1 likes14 downloads2y agoHugging Face28infinite-dataset-hub /ParagraphProof ParagraphProof tags: detection, paragraph, evidence-verification Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: This dataset contains paragraphs from various sources with the goal of identifying whether there is evidence supporting a specific claim within each paragraph. The dataset has been designed to aid in the development and training of machine learning models for the task of evidence-verification in textual data. Each… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphProof.tabularn<1K0 likes13 downloads2y agoHugging Face29infinite-dataset-hub /ParagraphSentinel ParagraphSentinel tags: Claim Analysis, Text Parsing, Anomaly Detection Note: This is an AI-generated dataset so its content may be inaccurate or false Dataset Description: The 'ParagraphSentinel' dataset is designed for the task of identifying and classifying claims within a text paragraph. Each entry in the dataset includes a paragraph of text and a label that indicates whether a claim is present and the type of claim identified. The dataset is suitable for machine learning models… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/ParagraphSentinel.tabularn<1K0 likes11 downloads2y agoHugging Face30mlfoundations-dev /d1_science_long_paragraphs_1k_eval_636d mlfoundations-dev/d1_science_long_paragraphs_1k_eval_636d Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces Accuracy 18.7 54.5 77.0 29.2 37.8 36.7 20.2 4.1 5.5 AIME24 Average Accuracy: 18.67% ± 1.26% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 23.33% 7 30 2 23.33% 7 30 3 23.33% 7 30 4… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/d1_science_long_paragraphs_1k_eval_636d.tabular1K<n<10K0 likes11 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.