CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01muhammedturan /turk-ictihat-kararlari-fulltext Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten) Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr) üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama sorguları ile toplanmıştır. Boyut Kayıt: 9,899,589 benzersiz karar Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052} Yıl aralığı: 1993-2026 Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor) Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.tabulartext-generation1M<n<10M1 likes555 downloads10d agoHugging Face02Madras1 /rag-qa-fulltext-ptbr RAG QA Full-Text PT-BR Mistral A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents using Mistral models. Every answer is anchored to literal quotations from the source text, making this dataset suitable for training and evaluating retrieval-augmented generation systems, extractive QA models, and reading comprehension benchmarks in Portuguese. Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.tabularquestion-answering1M<n<10M0 likes168 downloads5mo agoHugging Face03acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes165 downloads3y agoHugging Face04albertge /ni-20-clustered-fulltext-modernbert-sweep-20250107tabular10K<n<100K0 likes54 downloads2y agoHugging Face05SkyWhal3 /stxbp1-pubmed-central-fulltext source_datasets: - PubMed Central STXBP1 PubMed Central Full-Text Dataset v2 A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research. 🆕 Version 2 Updates (December 2025) Complete re-extraction with improved HTML parsing Full main text with proper section headers Enhanced metadata extraction 99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.tabulartext-generation10K<n<100K0 likes51 downloads9mo agoHugging Face06metehan777 /cc-aeo-geo-fulltext-CC-MAIN-2026-21tabular100K<n<1M0 likes50 downloads3mo agoHugging Face07rchu233 /ni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130tabular10K<n<100K0 likes48 downloads2y agoHugging Face08Mikimi /ru-wikipedia-daily-pageviews-full-texttabular10K<n<100K2 likes45 downloads9mo agoHugging Face09metehan777 /cc-turkish-fulltext-CC-MAIN-2026-21 CC-MAIN-2026-21 Turkish URLs 31.4M URLs · 538K domains from Common Crawl columnar index (content_languages contains tur, HTTP 200). Explorer Browse with pagination and domain search: CC Turkish Explorer Space Files File Description turkish_CC-MAIN-2026-21_urls.parquet 31.4M URLs (1 GB) turkish_CC-MAIN-2026-21_domain_leaderboard.parquet 538K domains turkish_CC-MAIN-2026-21_domain_leaderboard_webgraph.parquet + HC/PR/tier… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/cc-turkish-fulltext-CC-MAIN-2026-21.tabular10M<n<100M0 likes33 downloads3mo agoHugging Face10albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.0-20250109tabular10K<n<100K0 likes32 downloads2y agoHugging Face11albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.6-20250109tabular10K<n<100K0 likes32 downloads2y agoHugging Face12albertge /dolly-15k-clustered-fulltext-modernbert-sweep-20250106tabular10K<n<100K0 likes31 downloads2y agoHugging Face13albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.4-20250109tabular10K<n<100K0 likes31 downloads2y agoHugging Face14Mikimi /ru-wikipedia-100k-full-text-daily-stats-10-years 📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews **Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025** 📖 Описание Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет. Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.tabulartext-generation10K<n<100K1 likes27 downloads9mo agoHugging Face15acmc /beamit-annotated_full_texts_dataset Dataset Card for "beamit-annotated_full_texts_dataset" More Information needed tabular1K<n<10K0 likes25 downloads3y agoHugging Face16nibzard /narodne-novine-full-text-markdown Narodne Novine Full Text Markdown HTML-to-Markdown text extraction snapshot derived from the NN archive. Coverage Extracted acts: 96855 Failed acts: 157 Missing HTML embodiments: 150 Files texts.parquet failures.parquet metadata.json Notes Extraction prefers /hrv/printhtml, then falls back to /hrv/html. Conversion method: markitdown_html This snapshot does not mirror PDFs. tabulartext-generation10K<n<100K0 likes25 downloads6mo agoHugging Face17albertge /dolly-15k-clustered-fulltext-less-sweep-20250106tabular10K<n<100K0 likes24 downloads2y agoHugging Face18albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.8-20250109tabular10K<n<100K0 likes24 downloads2y agoHugging Face19albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p1.0-20250109tabular10K<n<100K0 likes24 downloads2y agoHugging Face20taesiri /ArXivSignals-FullText ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout A continuously-updated, day-partitioned dataset of arXiv papers converted to clean full text by a vision OCR pipeline: each paper's PDF is rendered to Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured layout JSON (typed, bounding-boxed blocks). It is the full-text companion to taesiri/ArXivSignals (metadata + LLM signal & summaries) and joins it on paper_id. How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.tabulartext-generation1K<n<10K0 likes23 downloads2mo agoHugging Face21albertge /dolly-15k-clustered-fulltext-agglomerative-16-20250103tabular10K<n<100K0 likes14 downloads2y agoHugging Face22AdrienB134 /Emilia-fr-tts-text-tags-full-v1tabular100K<n<1M0 likes11 downloads2y agoHugging Face23SF-Corpus /EF_Full_Texts Dataset Card for SF Nexus Extracted Features: Full Texts Dataset Summary The SF Nexus Extracted Features Full Texts dataset contains text and metadata from 403 mid-twentieth century science fiction books, originally digitized from Temple University Libraries' Paskow Science Fiction Collection. After digitization, the books were cleaned using Abbyy FineReader. Because this is a collection of copyrighted fiction, the books have been disaggregated. Each row of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/SF-Corpus/EF_Full_Texts.tabularn<1K1 likes10 downloads3y agoHugging Face24albertge /dolly-15k-clustered-fulltext-20250102tabular10K<n<100K0 likes10 downloads2y agoHugging Face25AdrienB134 /MOSEL-fr-tts-text-tags-full-v1tabular1M<n<10M0 likes9 downloads2y agoHugging Face26albertge /dolly-15k-clustered-less-fulltext-20250106tabular10K<n<100K0 likes9 downloads2y agoHugging Face27Titung /nepali-books-text-corpus-fullgatedtabular10K<n<100K0 likes7 downloads2mo agoHugging Face28AdrienB134 /fr-tts-text-tags-full-v1tabular100K<n<1M0 likes6 downloads2y agoHugging Face29amang1802 /synthetic_data_qna_fulltext_conditioned_L3.3_70Btabular10K<n<100K0 likes6 downloads2y agoHugging Face30Mikimi /ru-wikipedia-top-200k-full-texttabular10K<n<100K0 likes6 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.