datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.stxbp1-pubmed-central-fulltext
source_datasets:
- PubMed Central
STXBP1 PubMed Central Full-Text Dataset v2
A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research.
🆕 Version 2 Updates (December 2025)
Complete re-extraction with improved HTML parsing
Full main text with proper section headers
Enhanced metadata extraction
99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.ru-wikipedia-100k-full-text-daily-stats-10-years
📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews
**Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025**
📖 Описание
Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет.
Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.narodne-novine-full-text-markdown
Narodne Novine Full Text Markdown
HTML-to-Markdown text extraction snapshot derived from the NN archive.
Coverage
Extracted acts: 96855
Failed acts: 157
Missing HTML embodiments: 150
Files
texts.parquet
failures.parquet
metadata.json
Notes
Extraction prefers /hrv/printhtml, then falls back to /hrv/html.
Conversion method: markitdown_html
This snapshot does not mirror PDFs.
ArXivSignals-FullText
ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout
A continuously-updated, day-partitioned dataset of arXiv papers converted to
clean full text by a vision OCR pipeline: each paper's PDF is rendered to
Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured
layout JSON (typed, bounding-boxed blocks). It is the full-text companion to
taesiri/ArXivSignals
(metadata + LLM signal & summaries) and joins it on paper_id.
How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.
