datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.rag-qa-fulltext-ptbr
RAG QA Full-Text PT-BR Mistral
A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs
with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents
using Mistral models. Every answer is anchored to literal quotations from the
source text, making this dataset suitable for training and evaluating
retrieval-augmented generation systems, extractive QA models, and reading
comprehension benchmarks in Portuguese.
Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
ni-20-clustered-fulltext-modernbert-sweep-20250107stxbp1-pubmed-central-fulltext
source_datasets:
- PubMed Central
STXBP1 PubMed Central Full-Text Dataset v2
A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research.
🆕 Version 2 Updates (December 2025)
Complete re-extraction with improved HTML parsing
Full main text with proper section headers
Enhanced metadata extraction
99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.cc-aeo-geo-fulltext-CC-MAIN-2026-21ni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130ru-wikipedia-daily-pageviews-full-textcc-turkish-fulltext-CC-MAIN-2026-21
CC-MAIN-2026-21 Turkish URLs
31.4M URLs · 538K domains from Common Crawl columnar index (content_languages contains tur, HTTP 200).
Explorer
Browse with pagination and domain search:
CC Turkish Explorer Space
Files
File
Description
turkish_CC-MAIN-2026-21_urls.parquet
31.4M URLs (1 GB)
turkish_CC-MAIN-2026-21_domain_leaderboard.parquet
538K domains
turkish_CC-MAIN-2026-21_domain_leaderboard_webgraph.parquet
+ HC/PR/tier… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/cc-turkish-fulltext-CC-MAIN-2026-21.dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.0-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.6-20250109dolly-15k-clustered-fulltext-modernbert-sweep-20250106dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.4-20250109ru-wikipedia-100k-full-text-daily-stats-10-years
📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews
**Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025**
📖 Описание
Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет.
Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.beamit-annotated_full_texts_dataset
Dataset Card for "beamit-annotated_full_texts_dataset"
More Information needed
narodne-novine-full-text-markdown
Narodne Novine Full Text Markdown
HTML-to-Markdown text extraction snapshot derived from the NN archive.
Coverage
Extracted acts: 96855
Failed acts: 157
Missing HTML embodiments: 150
Files
texts.parquet
failures.parquet
metadata.json
Notes
Extraction prefers /hrv/printhtml, then falls back to /hrv/html.
Conversion method: markitdown_html
This snapshot does not mirror PDFs.
dolly-15k-clustered-fulltext-less-sweep-20250106dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.8-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p1.0-20250109ArXivSignals-FullText
ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout
A continuously-updated, day-partitioned dataset of arXiv papers converted to
clean full text by a vision OCR pipeline: each paper's PDF is rendered to
Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured
layout JSON (typed, bounding-boxed blocks). It is the full-text companion to
taesiri/ArXivSignals
(metadata + LLM signal & summaries) and joins it on paper_id.
How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.dolly-15k-clustered-fulltext-agglomerative-16-20250103Emilia-fr-tts-text-tags-full-v1EF_Full_Texts
Dataset Card for SF Nexus Extracted Features: Full Texts
Dataset Summary
The SF Nexus Extracted Features Full Texts dataset contains text and metadata from 403 mid-twentieth century science fiction books, originally digitized from Temple University Libraries' Paskow Science Fiction Collection.
After digitization, the books were cleaned using Abbyy FineReader.
Because this is a collection of copyrighted fiction, the books have been disaggregated.
Each row of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/SF-Corpus/EF_Full_Texts.dolly-15k-clustered-fulltext-20250102MOSEL-fr-tts-text-tags-full-v1dolly-15k-clustered-less-fulltext-20250106nepali-books-text-corpus-fullfr-tts-text-tags-full-v1synthetic_data_qna_fulltext_conditioned_L3.3_70Bru-wikipedia-top-200k-full-text
