datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beamit-full-texts-dataset
Dataset Card for "beamit-full-texts-dataset"
More Information needed
Gradients_Gradients_and_Text_Full_Logic_Captionsturk-ictihat-kararlari-fulltext
Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten)
Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr)
üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama
sorguları ile toplanmıştır.
Boyut
Kayıt: 9,899,589 benzersiz karar
Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052}
Yıl aralığı: 1993-2026
Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor)
Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.hamela_books_text_full_okarXiv-full-text-chunked
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.Pes2oX-fulltextIntroducing Pes2oX Full Text, a transformed dataset derived from the original Allen AI's Pes2o dataset. Our focus in this dataset was to restructure and reorganize the original Pes2o dataset. This was done to make it more accessible to research groups in terms of using it for training Artificial Intelligence models and fine-tuning for specific tasks within a particular domain.
Why was restructuring necessary?
After examining the original Pes2o dataset's structure, we found it necessary to… See the full description on the dataset page: https://huggingface.co/datasets/laion/Pes2oX-fulltext.shamela_books_text_full
Shamela_Books_Text_Full
This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis.
🔗 This dataset is linked to the companion metadata dataset:
👉 Shamela_Books_info via the book_id field.
Update :
The dataset includes the original raw files as well as a single… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/shamela_books_text_full.rag-textbook-instruct-full
Dataset Card for "rag-textbook-instruct-full"
More Information needed
rag-qa-fulltext-ptbr
RAG QA Full-Text PT-BR Mistral
A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs
with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents
using Mistral models. Every answer is anchored to literal quotations from the
source text, making this dataset suitable for training and evaluating
retrieval-augmented generation systems, extractive QA models, and reading
comprehension benchmarks in Portuguese.
Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.beamit-annotated-full-texts-dataset
Dataset Card for "beamit-annotated-full-texts-dataset"
More Information needed
stormfront-full-textonly
Dataset Card for "stormfront-full-textonly"
More Information needed
ewe-tts-bible-full-audio-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ewe Tts Bible Full Audio Text
alvarobartt-improving-text-embeddings-with-llms-full
Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml"
or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.cord-19-fulltext
Dataset Card for [pritamdeka/cord-19-fulltext]
Dataset Description
Dataset Summary
This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks.
Languages
English
Citation Information
@article{Wang2020CORD19TC,
title={CORD-19: The Covid-19 Open Research Dataset},
author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.arXiv-full-text-chunked-qaCurated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.task1292_yelp_review_full_text_categorization
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1292_yelp_review_full_text_categorization
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1292_yelp_review_full_text_categorization.text-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.license-plate-text-recognition-full
Dataset Card for "license-plate-text-recognition-full"
Background Information
This dataset is generated from keremberke/license-plate-object-detection dataset. What we have done is:
Get the Bounding Boxes for each plate in an image,
Crop the image to make the plate only visible,
Run it through the microsoft/trocr-large-printed model to extract the written information.
Structure of the Dataset
It has the same structure as the… See the full description on the dataset page: https://huggingface.co/datasets/sonnetechnology/license-plate-text-recognition-full.CeLLaTe-benchMark-set-fulltext
Dataset Card for CeLLaTe FullText Benchmark Dataset
Dataset summary
The CeLLaTe FullText Benchmark Dataset is a curated collection of biomedical fulltext created as a benchmark set for evaluating CeLLaTe named entity recognition (NER) models. It was extracted from Europe PMC article XML sources, with a focus on open-access articles. The dataset is intended to support model development and testing by providing an independent evaluation set for demo production runs… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/CeLLaTe-benchMark-set-fulltext.ni-20-clustered-fulltext-modernbert-sweep-20250107stxbp1-pubmed-central-fulltext
source_datasets:
- PubMed Central
STXBP1 PubMed Central Full-Text Dataset v2
A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research.
🆕 Version 2 Updates (December 2025)
Complete re-extraction with improved HTML parsing
Full main text with proper section headers
Enhanced metadata extraction
99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.cc-aeo-geo-fulltext-CC-MAIN-2026-21Instruction-text-only-fullsharegpt_formatted_cord19_fulltextni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130ru-wikipedia-daily-pageviews-full-textorpo-text-pairs-full
ORPO Text Preference Pairs (Full)
This dataset contains two versions of preference pairs for training language models using ORPO, DPO, or similar preference-based alignment methods.
Dataset Description
File
Rows
Description
orpo_pairs.jsonl
8,249
Refined/filtered pairs (recommended)
orpo_pairs_all.jsonl
14,214
Full dataset before filtering
Format: JSONL
Language: English
Task: Text-only preference learning (no images)
Schema
Each row… See the full description on the dataset page: https://huggingface.co/datasets/mncai/orpo-text-pairs-full.arXiv-full-text-chunked-testcc-turkish-fulltext-CC-MAIN-2026-21
CC-MAIN-2026-21 Turkish URLs
31.4M URLs · 538K domains from Common Crawl columnar index (content_languages contains tur, HTTP 200).
Explorer
Browse with pagination and domain search:
CC Turkish Explorer Space
Files
File
Description
turkish_CC-MAIN-2026-21_urls.parquet
31.4M URLs (1 GB)
turkish_CC-MAIN-2026-21_domain_leaderboard.parquet
538K domains
turkish_CC-MAIN-2026-21_domain_leaderboard_webgraph.parquet
+ HC/PR/tier… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/cc-turkish-fulltext-CC-MAIN-2026-21.
