CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3k downloads3y agoHugging Face02Felldude /Gradients_Gradients_and_Text_Full_Logic_Captionsimage1K<n<10K2 likes1.5k downloads13d agoHugging Face03muhammedturan /turk-ictihat-kararlari-fulltext Türk İçtihat Kararları - Tam Metin (UYAP e-Bedesten) Adalet Bakanlığı UYAP Mevzuat ve İçtihat Programı (bedesten.adalet.gov.tr) üzerinden derlenen Yargıtay kararları veri seti. Kavram bazlı arama sorguları ile toplanmıştır. Boyut Kayıt: 9,899,589 benzersiz karar Terim dağılımı: {"bosanma": 100052, "karar": 9894542, "nafaka": 66241, "tazminat": 100052} Yıl aralığı: 1993-2026 Tam metni gelen karar: 39,301 (artımlı olarak dolduruluyor) Şema… See the full description on the dataset page: https://huggingface.co/datasets/muhammedturan/turk-ictihat-kararlari-fulltext.tabulartext-generation1M<n<10M1 likes555 downloads10d agoHugging Face04MathematicianNLPer /hamela_books_text_full_oktext1M<n<10M1 likes501 downloads7mo agoHugging Face05amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes453 downloads2y agoHugging Face06laion /Pes2oX-fulltextIntroducing Pes2oX Full Text, a transformed dataset derived from the original Allen AI's Pes2o dataset. Our focus in this dataset was to restructure and reorganize the original Pes2o dataset. This was done to make it more accessible to research groups in terms of using it for training Artificial Intelligence models and fine-tuning for specific tasks within a particular domain. Why was restructuring necessary? After examining the original Pes2o dataset's structure, we found it necessary to… See the full description on the dataset page: https://huggingface.co/datasets/laion/Pes2oX-fulltext.text1M<n<10M1 likes182 downloads2y agoHugging Face07MoMonir /shamela_books_text_full Shamela_Books_Text_Full This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis. 🔗 This dataset is linked to the companion metadata dataset: 👉 Shamela_Books_info via the book_id field. Update : The dataset includes the original raw files as well as a single… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/shamela_books_text_full.text10M<n<100M2 likes171 downloads4mo agoHugging Face08open-phi /rag-textbook-instruct-full Dataset Card for "rag-textbook-instruct-full" More Information needed text1K<n<10K10 likes168 downloads3y agoHugging Face09Madras1 /rag-qa-fulltext-ptbr RAG QA Full-Text PT-BR Mistral A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents using Mistral models. Every answer is anchored to literal quotations from the source text, making this dataset suitable for training and evaluating retrieval-augmented generation systems, extractive QA models, and reading comprehension benchmarks in Portuguese. Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.tabularquestion-answering1M<n<10M0 likes168 downloads5mo agoHugging Face10acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes165 downloads3y agoHugging Face11kastan /stormfront-full-textonly Dataset Card for "stormfront-full-textonly" More Information needed text10M<n<100M0 likes162 downloads4y agoHugging Face12ghananlpcommunity /ewe-tts-bible-full-audio-text This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ewe Tts Bible Full Audio Text audio1K<n<10K0 likes152 downloads3mo agoHugging Face13distilabel-internal-testing /alvarobartt-improving-text-embeddings-with-llms-full Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.textn<1K0 likes142 downloads2y agoHugging Face14pritamdeka /cord-19-fulltext Dataset Card for [pritamdeka/cord-19-fulltext] Dataset Description Dataset Summary This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks. Languages English Citation Information @article{Wang2020CORD19TC, title={CORD-19: The Covid-19 Open Research Dataset}, author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.text100K<n<1M2 likes106 downloads5y agoHugging Face15amrachraf /arXiv-full-text-chunked-qatext100K<n<1M0 likes103 downloads2y agoHugging Face16crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes89 downloads1y agoHugging Face17Lots-of-LoRAs /task1292_yelp_review_full_text_categorization Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1292_yelp_review_full_text_categorization Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1292_yelp_review_full_text_categorization.texttext-generation1K<n<10K0 likes86 downloads2y agoHugging Face18datapointai /text-2-image-dpo-human-preferences-fullgated Text-2-Image DPO Human Preferences (Full) The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference. This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see: datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.imageimage-classification10K<n<100K1 likes86 downloads6mo agoHugging Face19sonnetechnology /license-plate-text-recognition-full Dataset Card for "license-plate-text-recognition-full" Background Information This dataset is generated from keremberke/license-plate-object-detection dataset. What we have done is: Get the Bounding Boxes for each plate in an image, Crop the image to make the plate only visible, Run it through the microsoft/trocr-large-printed model to extract the written information. Structure of the Dataset It has the same structure as the… See the full description on the dataset page: https://huggingface.co/datasets/sonnetechnology/license-plate-text-recognition-full.imageimage-to-text1K<n<10K3 likes73 downloads3y agoHugging Face20OTAR3088 /CeLLaTe-benchMark-set-fulltext Dataset Card for CeLLaTe FullText Benchmark Dataset Dataset summary The CeLLaTe FullText Benchmark Dataset is a curated collection of biomedical fulltext created as a benchmark set for evaluating CeLLaTe named entity recognition (NER) models. It was extracted from Europe PMC article XML sources, with a focus on open-access articles. The dataset is intended to support model development and testing by providing an independent evaluation set for demo production runs… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/CeLLaTe-benchMark-set-fulltext.text10K<n<100K0 likes65 downloads1mo agoHugging Face21albertge /ni-20-clustered-fulltext-modernbert-sweep-20250107tabular10K<n<100K0 likes54 downloads2y agoHugging Face22SkyWhal3 /stxbp1-pubmed-central-fulltext source_datasets: - PubMed Central STXBP1 PubMed Central Full-Text Dataset v2 A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research. 🆕 Version 2 Updates (December 2025) Complete re-extraction with improved HTML parsing Full main text with proper section headers Enhanced metadata extraction 99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.tabulartext-generation10K<n<100K0 likes51 downloads9mo agoHugging Face23metehan777 /cc-aeo-geo-fulltext-CC-MAIN-2026-21tabular100K<n<1M0 likes50 downloads3mo agoHugging Face24Menlo /Instruction-text-only-fulltext10K<n<100K3 likes48 downloads1y agoHugging Face25PranavHarshan /sharegpt_formatted_cord19_fulltexttext100K<n<1M0 likes48 downloads2y agoHugging Face26rchu233 /ni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130tabular10K<n<100K0 likes48 downloads2y agoHugging Face27Mikimi /ru-wikipedia-daily-pageviews-full-texttabular10K<n<100K2 likes45 downloads9mo agoHugging Face28mncai /orpo-text-pairs-full ORPO Text Preference Pairs (Full) This dataset contains two versions of preference pairs for training language models using ORPO, DPO, or similar preference-based alignment methods. Dataset Description File Rows Description orpo_pairs.jsonl 8,249 Refined/filtered pairs (recommended) orpo_pairs_all.jsonl 14,214 Full dataset before filtering Format: JSONL Language: English Task: Text-only preference learning (no images) Schema Each row… See the full description on the dataset page: https://huggingface.co/datasets/mncai/orpo-text-pairs-full.text10K<n<100K0 likes38 downloads8mo agoHugging Face29amrachraf /arXiv-full-text-chunked-testtextn<1K0 likes36 downloads2y agoHugging Face30metehan777 /cc-turkish-fulltext-CC-MAIN-2026-21 CC-MAIN-2026-21 Turkish URLs 31.4M URLs · 538K domains from Common Crawl columnar index (content_languages contains tur, HTTP 200). Explorer Browse with pagination and domain search: CC Turkish Explorer Space Files File Description turkish_CC-MAIN-2026-21_urls.parquet 31.4M URLs (1 GB) turkish_CC-MAIN-2026-21_domain_leaderboard.parquet 538K domains turkish_CC-MAIN-2026-21_domain_leaderboard_webgraph.parquet + HC/PR/tier… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/cc-turkish-fulltext-CC-MAIN-2026-21.tabular10M<n<100M0 likes33 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.