datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.arXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.task1292_yelp_review_full_text_categorization
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1292_yelp_review_full_text_categorization
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1292_yelp_review_full_text_categorization.stxbp1-pubmed-central-fulltext
source_datasets:
- PubMed Central
STXBP1 PubMed Central Full-Text Dataset v2
A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research.
🆕 Version 2 Updates (December 2025)
Complete re-extraction with improved HTML parsing
Full main text with proper section headers
Enhanced metadata extraction
99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.ru-wikipedia-100k-full-text-daily-stats-10-years
📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews
**Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025**
📖 Описание
Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет.
Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.narodne-novine-full-text-markdown
Narodne Novine Full Text Markdown
HTML-to-Markdown text extraction snapshot derived from the NN archive.
Coverage
Extracted acts: 96855
Failed acts: 157
Missing HTML embodiments: 150
Files
texts.parquet
failures.parquet
metadata.json
Notes
Extraction prefers /hrv/printhtml, then falls back to /hrv/html.
Conversion method: markitdown_html
This snapshot does not mirror PDFs.
ArXivSignals-FullText
ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout
A continuously-updated, day-partitioned dataset of arXiv papers converted to
clean full text by a vision OCR pipeline: each paper's PDF is rendered to
Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured
layout JSON (typed, bounding-boxed blocks). It is the full-text companion to
taesiri/ArXivSignals
(metadata + LLM signal & summaries) and joins it on paper_id.
How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.
