CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /ghana-sentences Ghana Sentences A growing sentence-level text corpus for Ghanaian languages, tagged with ISO 639-3 codes and split into per-language subsets. The goal is broad-coverage text across all Ghanaian languages; this first release draws on school curriculum materials and a licensing-exam benchmark. More sources will be added over time. Language list and ISO codes follow GhanaNLP/ghana-taught-local-languages. Loading from datasets import load_dataset everything =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-sentences.texttext-generation100K<n<1M0 likes704 downloads2mo agoHugging Face02agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes316 downloads2y agoHugging Face03hmar-heritage-org /sentencesgated sentences A multi-register sentence corpus for the Hmar language (hmr, ISO 639-3), containing 260,492 train sentences and 5,316 evaluation sentences (~4.15 million words). Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev). Overview Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241) Family: Zo Languages Volume: 260,492 train sentences | 5,316 evaluation sentences (265,808 total, ~4.15M words) Validation: 100% verified via hmaraniam… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/sentences.textfill-mask100K<n<1M2 likes207 downloads6d agoHugging Face04neurlang /low-quality-multilingual-sentences Low Quality Multilingual Sentences This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages. The new sentences in this dataset are low quality, proceed with caution. texttext-generation1K<n<10K1 likes198 downloads5mo agoHugging Face05marcolecci /Chinese_sentencestext100K<n<1M0 likes158 downloads3y agoHugging Face06a3xrfgb /amharic-sentences-corpus Amharic Sentences Corpus V1.0 Source: Telegram This 1.6 million Amharic sentences corpus reflects current Amharic usage as of December 20, 2025, and is designed for anyone interested in: Training Amharic-based LLMs Fine-tuning NLP models Building search, summarization, or generative systems in Amharic The dataset is heavily cleaned and normalized, but like any serious LLM dataset, it still needs proper tokenization for pre-training. I recommend using an… See the full description on the dataset page: https://huggingface.co/datasets/a3xrfgb/amharic-sentences-corpus.text1M<n<10M1 likes158 downloads7mo agoHugging Face07danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes152 downloads10mo agoHugging Face08danielrosehill /English-Hebrew-Mixed-Sentences English-Hebrew Mixed Sentences Dataset A dataset of English sentences with Hebrew words and phrases interspersed, designed for speech-to-text training and evaluation for English speakers in Israel. Overview This dataset addresses a common challenge for English-speaking immigrants in Israel: standard speech-to-text (STT) systems struggle to accurately transcribe code-switched speech where Hebrew words are mixed into primarily English sentences. Example: "I need to pick up… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/English-Hebrew-Mixed-Sentences.audion<1K0 likes63 downloads10mo agoHugging Face09ugursa /Yahoo-Finance-News-Sentences Yahoo Finance News Sentences Dataset This dataset was created from Yahoo Finance News Articles collected from 6.12.2023 to 20.12.2023 texttext-classification10K<n<100K18 likes59 downloads3y agoHugging Face10murat /kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters Kyrgyz Orthographic Correction Dataset Dataset Description This dataset is designed to fine-tune language models for a Kyrgyz-to-Kyrgyz orthographic correction task. It addresses a common issue in digital Kyrgyz text where specific Cyrillic characters specific to the Kyrgyz language (ө, ң, ү) are replaced by their Russian keyboard counterparts (о, н, у). The dataset is structured in a conversational format, making it ideal for instruction-tuning chat models.… See the full description on the dataset page: https://huggingface.co/datasets/murat/kyrgyz_sentences_with_incorrect_and_correct_umlaut_characters.text1M<n<10M0 likes59 downloads1y agoHugging Face11drewoodward /spanglish-sentences Spanglish Sentences A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models. Data format Each line of spanglish_sentences.jsonl is a JSON object with two fields: field description sentence A Spanglish utterance (mixed Spanish / English, or monolingual in either language). english_translation The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.texttranslation10K<n<100K0 likes51 downloads5mo agoHugging Face12wrootpress /bonaventure-sentences Bonaventure on the Sentences (Latin ↔ English) 2,113 chunks of Bonaventure's Commentary on the Sentences and related works, Latin (Quaracchi) aligned with English (~3.33M words). Apparatus (notes, scholia) is kept in a separate field, never merged into the text. Canonical home: https://bonaventure.wrootpress.com (each record carries its canonical URL). This dataset is a machine-generated export of that site's build, regenerated from the source of truth and never hand-edited; the… See the full description on the dataset page: https://huggingface.co/datasets/wrootpress/bonaventure-sentences.text1K<n<10K0 likes42 downloads6d agoHugging Face13librarian-bots /model-card-sentences-annotatedtexttoken-classification100K<n<1M4 likes39 downloads3y agoHugging Face14gplsi /alia_multilingual_parallel_sentences MULTILINGUAL PARALLEL SENTENCES Dataset The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models. It provides aligned sentences in multiple languages to facilitate multilingual learning. Dataset Structure The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language. The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.texttext-generation1M<n<10M0 likes36 downloads8mo agoHugging Face15ajsbsd /legalese-sentences_estonian-english 📚 Estonian → English Legal Sentence Translation Dataset @ajsbsd | Non-commercial Use Only This dataset contains Estonian legal sentences translated into English, derived from the original paulpall/legalese-sentences_estonian dataset. Translations were generated using the Helsinki-NLP/opus-mt-et-en model. 🛥️ Dedication This dataset is dedicated to the pursuit of truth, transparency, and informed discourse in a world shaped by complex global threats.… See the full description on the dataset page: https://huggingface.co/datasets/ajsbsd/legalese-sentences_estonian-english.text10K<n<100K1 likes33 downloads1y agoHugging Face16harish03 /english_hinglist_sentencestext100K<n<1M1 likes31 downloads3y agoHugging Face17Reubencf /Adaption-multilingual-sentences This dataset is a remastered version of Reubencf/PolyglotText prepared using Adaption's Adaptive Data platform. Multilingual Sentences (Adaption) 9,999 sentences across 123 languages. A broad multilingual subset of PolyglotText — originally derived from the Tatoeba project — with Adaption-sharpened enhanced_prompt / enhanced_completion / reasoning_trace columns. Each row carries a source-language sentence, translations, and the Adaption-processed fields. Dataset size… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-sentences.texttranslation1K<n<10K1 likes31 downloads5mo agoHugging Face18toni5rovic /bcms-claim-sentencestexttext-classification10K<n<100K0 likes24 downloads1y agoHugging Face19mostafaamiri /fa-topic-sentences README for fa-topic-sentences Dataset Overview The fa-topic-sentences dataset is a comprehensive collection of sentences categorized into various topics. Each topic contains approximately 50 sentences in Persian, accompanied by a paraphrased version of each sentence. The dataset is structured in JSON format, providing a straightforward method for accessing individual entries. Topics Included The dataset encompasses the following topics: History Fashion… See the full description on the dataset page: https://huggingface.co/datasets/mostafaamiri/fa-topic-sentences.textn<1K0 likes23 downloads2y agoHugging Face20jfernandez /cebuano-filipino-sentencestext100K<n<1M4 likes22 downloads4y agoHugging Face21kaengreg /eng-xquad-sentencestext1K<n<10K0 likes22 downloads2y agoHugging Face22kaengreg /rus-xquad-sentencestext1K<n<10K0 likes20 downloads2y agoHugging Face23esat-krky /processed-turkish-sentencestext1M<n<10M0 likes16 downloads2mo agoHugging Face24nadizik /tech-sentences-error-robustness Technical English Sentences for NLP Robustness (50k) Dataset Description This dataset contains 50,000 synthetically generated English sentences tailored around Software Engineering, DevOps, and IT contexts. Key Feature: Intentional Grammatical Anomalies A unique characteristic of this dataset is the presence of intentional grammatical and morphological anomalies in verb forms (e.g., combining past tense with third-person singular endings like… See the full description on the dataset page: https://huggingface.co/datasets/nadizik/tech-sentences-error-robustness.texttext-classification10K<n<100K0 likes15 downloads4mo agoHugging Face25readallaboutit /tbc-sentencestext100M<n<1B0 likes14 downloads1y agoHugging Face26Reubencf /multilingual-single-sentences This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. multilingual_single_sentences This dataset consists of single-sentence completions spanning a diverse range of languages, including Russian, German, Spanish, Korean, Chinese, English, French, and Japanese. The content varies widely, covering topics from historical restoration and travel logistics to proverbs and daily observations. Each entry is presented as an isolated text… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-single-sentences.tabular10K<n<100K0 likes13 downloads5mo agoHugging Face27agentlans /expanded-english-sentences Expanded English Sentences Dataset This dataset includes over 15 000 random sentences from the agentlans/high-quality-english-sentences dataset, each paired with a paragraph generated by a customized Llama 3.1 8B model, providing additional context. Overview train.jsonl.gz: Contains original sentences and their corresponding AI-generated paragraphs in JSONL (JSON Lines) format compressed using GZip. Variable Definition Type sentence Original sentence from the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/expanded-english-sentences.texttext-generation10K<n<100K1 likes11 downloads2y agoHugging Face28Reubencf /marathi-czech-sentences This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. marathi_czech_sentences This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures. Dataset size There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.tabular1K<n<10K0 likes11 downloads5mo agoHugging Face29Uyghur-Corpus /uyghur-sentencesgated 🌟 Uyghur AI Corpus: Bridging Heritage & Technology 🌟 ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى 🌹 Introduction / كىرىش سۆز In the era of Artificial Intelligence, language is data, and data is survival.&nbsp; The Uyghur AI Corpus is an initiative to ensure the Uyghur language thrives in the digital age. This dataset serves as a foundational resource to train Large Language Models (LLMs), enabling them to understand, generate… See the full description on the dataset page: https://huggingface.co/datasets/Uyghur-Corpus/uyghur-sentences.texttext-generation100K<n<1M0 likes11 downloads2mo agoHugging Face30mideind /icelandic-sentences-gectextn<1K0 likes10 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.