CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ghanaopenai /ghana-sentences Ghana Sentences A growing sentence-level text corpus for Ghanaian languages, tagged with ISO 639-3 codes and split into per-language subsets. The goal is broad-coverage text across all Ghanaian languages; this first release draws on school curriculum materials and a licensing-exam benchmark. More sources will be added over time. Language list and ISO codes follow GhanaNLP/ghana-taught-local-languages. Loading from datasets import load_dataset everything =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-sentences.texttext-generation100K<n<1M0 likes704 downloads2mo agoHugging Face02agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes316 downloads2y agoHugging Face03neurlang /low-quality-multilingual-sentences Low Quality Multilingual Sentences This dataset is a complement to agentlans/high-quality-multilingual-sentences to extend it to more languages. The new sentences in this dataset are low quality, proceed with caution. texttext-generation1K<n<10K1 likes198 downloads5mo agoHugging Face04drewoodward /spanglish-sentences Spanglish Sentences A dataset of 10,576 Spanish–English code-switched ("Spanglish") sentences paired with English translations, intended for training and evaluating code-switch translation models. Data format Each line of spanglish_sentences.jsonl is a JSON object with two fields: field description sentence A Spanglish utterance (mixed Spanish / English, or monolingual in either language). english_translation The English translation. When the source is… See the full description on the dataset page: https://huggingface.co/datasets/drewoodward/spanglish-sentences.texttranslation10K<n<100K0 likes51 downloads5mo agoHugging Face05gplsi /alia_multilingual_parallel_sentences MULTILINGUAL PARALLEL SENTENCES Dataset The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models. It provides aligned sentences in multiple languages to facilitate multilingual learning. Dataset Structure The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language. The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.texttext-generation1M<n<10M0 likes36 downloads8mo agoHugging Face06Reubencf /Adaption-multilingual-sentences This dataset is a remastered version of Reubencf/PolyglotText prepared using Adaption's Adaptive Data platform. Multilingual Sentences (Adaption) 9,999 sentences across 123 languages. A broad multilingual subset of PolyglotText — originally derived from the Tatoeba project — with Adaption-sharpened enhanced_prompt / enhanced_completion / reasoning_trace columns. Each row carries a source-language sentence, translations, and the Adaption-processed fields. Dataset size… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/Adaption-multilingual-sentences.texttranslation1K<n<10K1 likes31 downloads5mo agoHugging Face07agentlans /expanded-english-sentences Expanded English Sentences Dataset This dataset includes over 15 000 random sentences from the agentlans/high-quality-english-sentences dataset, each paired with a paragraph generated by a customized Llama 3.1 8B model, providing additional context. Overview train.jsonl.gz: Contains original sentences and their corresponding AI-generated paragraphs in JSONL (JSON Lines) format compressed using GZip. Variable Definition Type sentence Original sentence from the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/expanded-english-sentences.texttext-generation10K<n<100K1 likes11 downloads2y agoHugging Face08Uyghur-Corpus /uyghur-sentencesgated 🌟 Uyghur AI Corpus: Bridging Heritage & Technology 🌟 ئۇيغۇرچە سۈنئىي ئىدراك خەزىنىسى: مىراس ۋە تېخنىكا كۆۋرۈكى 🌹 Introduction / كىرىش سۆز In the era of Artificial Intelligence, language is data, and data is survival.&nbsp; The Uyghur AI Corpus is an initiative to ensure the Uyghur language thrives in the digital age. This dataset serves as a foundational resource to train Large Language Models (LLMs), enabling them to understand, generate… See the full description on the dataset page: https://huggingface.co/datasets/Uyghur-Corpus/uyghur-sentences.texttext-generation100K<n<1M0 likes11 downloads2mo agoHugging Face09agentlans /Taiwan-Text-Excellence-sentencesgated 台灣文摘句資料集 概述 Taiwan-Text-Excellence 句子資料集是從較大的 liswei/Taiwan-Text-Excellence-2B 資料集中抽取的 200 萬個獨特中文句子的綜合集。這些句子是隨機選取的,並使用 chinese-sentence-processor 工具進行分割。此資料集非常適合各種自然語言處理任務,包括語言建模、文本生成和其他研究用途。 資料集統計 總句數: 2,000,000 訓練集: 1,600,000 個句子 測試集: 400,000 個句子 資料格式 資料集中的每一行都包含一個欄位: **text**:包含中文句子的字串。 範例 {"text": "而新郎和女方家人的脂燭在當晚亦會合二為一,再送到母屋帳前點燃一個燈籠,保持三日不滅。"} {"text": "這個時期簽署的現代劇至今仍是台灣戲劇的中堅力量,而這十年為之奮鬥也奠定了其後數十年的基礎。"} {"text": "曾柏瑜今天也車票,明天還請在規劃畫相關票活動中。"}… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/Taiwan-Text-Excellence-sentences.texttext-generation1M<n<10M0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.