CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AnimaLab /bias-test-gpt-sentences Dataset Card for "BiasTestGPT: Generated Test Sentences" Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models. This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool. BiasTestGPT HuggingFace Tool Dataset with Bias Specifications Project Landing Page Dataset Structure The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.text1K<n<10K1 likes920 downloads3y agoHugging Face02Anon3365 /bias-test-gpt-sentencestext1K<n<10K0 likes303 downloads3y agoHugging Face03dvgodoy /yoda_sentences Yoda Speak This small dataset was built using two resources: Harvard Sentences, a list of 720 short sentences grouped into 72 sets of 10 sentences each English to Yoda Translator, an online translator that converts normal English into Yoda's way of speaking. Fun with this dataset I hope you have! Yes, hrrrm. texttranslationn<1K8 likes267 downloads2y agoHugging Face04TechWolf /Synthetic-ESCO-skill-sentences Synthetic job ads for all ESCO skills Dataset Summary This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0. Languages We use the English version of ESCO, and all generated sentences are in English. Dataset Structure The dataset consists of 138,260 (sentence, skill) pairs. Citation Information If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.texttext-classification100K<n<1M17 likes165 downloads2y agoHugging Face05cassiehu /probing_sentences_liwctextn<1K0 likes165 downloads2y agoHugging Face06paulpall /textbooks-sentences_estonian Corpus for Learners of Estonian as a Second Language 2022 with Synthetic Grammatical Errors Dataset Summary The "Corpus for Learners of Estonian as a Second Language 2022" (Eesti keele kui teise keele õppekorpus 2022) is a specialized linguistic resource designed to support learners of Estonian as a second language. The corpus is composed of sentences extracted from 34 different Estonian as a Second Language coursebooks, ranging from A1 to C1 levels. We have introduced… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/textbooks-sentences_estonian.texttranslation10K<n<100K0 likes119 downloads2y agoHugging Face07cpllab /syntaxgym_sentencestabular1K<n<10K1 likes107 downloads4y agoHugging Face08stjiris /portuguese-legal-sentences-v0 Work developed as part of Project IRIS. Thesis: A Semantic Search System for Supremo Tribunal de Justiça Portuguese Legal Sentences Collection of Legal Sentences from the Portuguese Supreme Court of Justice The goal of this dataset was to be used for MLM and TSDAE Contributions @rufimelo99 If you use this work, please cite: @InProceedings{MeloSemantic, author="Melo, Rui and Santos, Pedro A. and Dias, Jo{\~a}o", editor="Moniz, Nuno and Vale, Zita and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.text1M<n<10M14 likes100 downloads2y agoHugging Face09duwuonline /en_vi_advanced_sentences Model description This data I crawled from these site: https://prep.vn/blog/idiom-theo-chu-de-trong-tieng-anh/ and https://www.enewsdispatch.com/ Idiom site I carefully translation, however, the enews site I use google translate texttranslationn<1K2 likes73 downloads3y agoHugging Face10Mina-Rajaei-Moghadam /US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.tabulartext-classification10K<n<100K2 likes41 downloads10mo agoHugging Face11erickfmm /agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs The code for processing can be found here Useful for data distillation, training or benchmarking. Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.tabularsentence-similarity1M<n<10M0 likes40 downloads1y agoHugging Face12SunayanaGawde /admin-test-en-mr-kon-parallel-500-sentencestexttranslationn<1K0 likes40 downloads1mo agoHugging Face13paulpall /Tartu-L2-sentences_estonian Tartu-L2 Corpus Dataset Summary The Tartu-L2 corpus is a comprehensive dataset designed for Estonian Grammatical Error Correction (GEC) research. Developed at Tartu University, it is the oldest and largest corpus in the domain. The corpus was created in two phases: 2004-2006 and 2018-2019, funded by the Estonian National Programme for Language Technology. Corpus Creation The initiative and structure were developed by Heiki-Jaan Kaalep. The corpus was originally… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/Tartu-L2-sentences_estonian.texttranslation1K<n<10K0 likes35 downloads2y agoHugging Face14cassiehu /probing_sentences_liwc_2textn<1K0 likes33 downloads2y agoHugging Face15marmolpen3 /slas-obligations-rights-sentencestexttext-classificationn<1K1 likes31 downloads3y agoHugging Face16finnstrom3693 /leipzig_en_simple_wikipedia_2021_sentences_100ktext100K<n<1M0 likes31 downloads2y agoHugging Face17paulpall /legalese-sentences_estonian Estonian Legalese Corpus Dataset Summary The Estonian Legal Texts dataset is a collection of legal documents extracted from the Estonian National Corpus. It is tailored for Natural Language Processing (NLP) tasks, particularly those involving the Estonian language. The dataset contains legal texts such as acts, regulations, and legal proceedings, making it an essential resource for developing language models, text classification systems, and other NLP tools for Estonian.… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/legalese-sentences_estonian.texttranslation10K<n<100K1 likes30 downloads2y agoHugging Face18shahxeebhassan /human_vs_ai_sentences Dataset Description This dataset contains 105,000 sentences, each labeled as either human-written (0) or AI-generated (1). It is designed for text classification tasks, particularly for distinguishing between human and AI-generated text. Dataset Structure Number of Instances: 105,000 sentences Labels: 0: Human-written 1: AI-generated Usage This dataset can be used to train models for text classification tasks. Below is an example of how to load and use the… See the full description on the dataset page: https://huggingface.co/datasets/shahxeebhassan/human_vs_ai_sentences.texttext-classification100K<n<1M10 likes28 downloads2y agoHugging Face19jaimevera1107 /similarity-sentences-spanish similarity-sentences-spanish (SSS) Dataset Summary This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics. The dataset also includes sentences from two existing datasets, STS-ES and STSB-Multi-MT, as well as SICK, which were used as additional sources. The sentences in this dataset were generated to exhibit varying levels of similarity based on randomly divided prompts. Source Share (rows) Count (rows) Score… See the full description on the dataset page: https://huggingface.co/datasets/jaimevera1107/similarity-sentences-spanish.textsentence-similarity10K<n<100K7 likes26 downloads3y agoHugging Face20agentlans /finewebedu-sentences Fineweb-edu Sentences Description: A dataset of sentences collected from the web. The dataset was created by splitting the text into individual sentences using the spaCy package, then removing duplicates and filtering for complete sentences in a semi-automated process. Source: HuggingFaceFW/fineweb-edu Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer. Annotations: The source field contains the URL of each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.text100K<n<1M0 likes25 downloads2y agoHugging Face21amedcj /10452_kurmanji-corrected-sentences Cleaned Kurmanji Kurdish Sentences Dataset (Hawar Standard) Dataset Description This dataset contains over 10,000 highly curated and grammatically corrected Kurmanji Kurdish sentences. While the original raw sentences were sourced from the open-source Tatoeba project, they have undergone extensive and meticulous editorial correction to meet the strict standards of the Hawar orthography and authentic Kurdish grammar (Celadet Alî Bedirxan rules). The… See the full description on the dataset page: https://huggingface.co/datasets/amedcj/10452_kurmanji-corrected-sentences.texttext-generation10K<n<100K0 likes22 downloads1mo agoHugging Face22orisuchy /Descriptive_Sentences_Hetext1K<n<10K2 likes20 downloads5y agoHugging Face23AnnikaSimonsen /GPT-4_FO-EN_parallel_blog_sentences_MQMThis is dataset contains 425 Faroese-to-English parallel sentences generated by GPT-4 that have been annotated by a single native speaker of Faroese using the Multidimensional Quality Metrics framework (MQM). The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus. In addition to the parallel sentences and human evaluation, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences_MQM.textn<1K0 likes19 downloads2y agoHugging Face24AnnikaSimonsen /GPT-4_FO-EN_parallel_blog_sentencesThis is dataset contains 1,673 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus. In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article. Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences.text1K<n<10K0 likes17 downloads2y agoHugging Face25AnnikaSimonsen /GPT-4_FO-EN_parallel_news_sentencesThis is dataset contains 3,735 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is news text from the Basic Language Resource Kit for Faroese 1.0 text corpus. In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article. Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_news_sentences.text1K<n<10K0 likes17 downloads2y agoHugging Face26paulpall /Tallinn-L2-sentences_estonian Tallinn-L2 Corpus Dataset Summary The Tallinn-L2 corpus is a significant dataset developed by the Language Technology Research Group at Tallinn University. This corpus is a valuable resource for Grammatical Error Correction (GEC) research, particularly in the Estonian language. The dataset consists of 3,790 sentences annotated in the Max-Match (M2) format, highlighting the type and location of errors made by learners of Estonian. Features Total Sentences:… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/Tallinn-L2-sentences_estonian.texttranslation1K<n<10K0 likes17 downloads2y agoHugging Face27KnoxDevelopers /english_maasai_pair_sentences English_Maasai_dataset 31,103 pair sentences (verses) extracted from the English and Maasai Bibles. This dataset has not been verified by any person who knows and understands both English & Maa languages. This dataset has not been cleaned very well and may still contain some noise. texttranslation10K<n<100K0 likes17 downloads2mo agoHugging Face28ambrosfitz /synth_history_sentences Synthetically generated history text, segemented into sentences. text1K<n<10K1 likes16 downloads2y agoHugging Face29KnoxDevelopers /english_kalenjin_pair_sentences English_Kalenjin_dataset 31,097 pair sentences (verses) extracted from the English and Kalenjin Bibles. This dataset has not been verified by any person who knows and understands both English & Kalenjin languages. This dataset has not been cleaned very well and may still contain some noise. text10K<n<100K0 likes16 downloads2mo agoHugging Face30adambuttrick /360K-funding-statement-sentences-name-identifiertext10M<n<100M1 likes15 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.