datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bias-test-gpt-sentences
Dataset Card for "BiasTestGPT: Generated Test Sentences"
Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models.
This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool.
BiasTestGPT HuggingFace Tool
Dataset with Bias Specifications
Project Landing Page
Dataset Structure
The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.bias-test-gpt-sentencesyoda_sentences
Yoda Speak
This small dataset was built using two resources:
Harvard Sentences, a list of 720 short sentences grouped into 72 sets of 10 sentences each
English to Yoda Translator, an online translator that converts normal English into Yoda's way of speaking.
Fun with this dataset I hope you have! Yes, hrrrm.
Synthetic-ESCO-skill-sentences
Synthetic job ads for all ESCO skills
Dataset Summary
This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0.
Languages
We use the English version of ESCO, and all generated sentences are in English.
Dataset Structure
The dataset consists of 138,260 (sentence, skill) pairs.
Citation Information
If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.probing_sentences_liwctextbooks-sentences_estonian
Corpus for Learners of Estonian as a Second Language 2022 with Synthetic Grammatical Errors
Dataset Summary
The "Corpus for Learners of Estonian as a Second Language 2022" (Eesti keele kui teise keele õppekorpus 2022) is a specialized linguistic resource designed to support learners of Estonian as a second language. The corpus is composed of sentences extracted from 34 different Estonian as a Second Language coursebooks, ranging from A1 to C1 levels. We have introduced… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/textbooks-sentences_estonian.syntaxgym_sentencesportuguese-legal-sentences-v0
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Contributions
@rufimelo99
If you use this work, please cite:
@InProceedings{MeloSemantic,
author="Melo, Rui
and Santos, Pedro A.
and Dias, Jo{\~a}o",
editor="Moniz, Nuno
and Vale, Zita
and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.en_vi_advanced_sentences
Model description
This data I crawled from these site: https://prep.vn/blog/idiom-theo-chu-de-trong-tieng-anh/ and https://www.enewsdispatch.com/
Idiom site I carefully translation, however, the enews site I use google translate
US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences
We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.agentlans__multilingual-sentences__paired_10_stsSentences from agentlans/multilingual-sentences in Spanish, and processed with Sentence Similarity Cosine Scores with model hiiamsid/sentence_similarity_spanish_es
Each sentence in original dataset was randomly assigned 10 rows (sentences) within a batch of 1000, calculate the sentence similarity, and then deleted duplicate pairs
The code for processing can be found here
Useful for data distillation, training or benchmarking.
Its recommended resampling the dataset to undersample to get a… See the full description on the dataset page: https://huggingface.co/datasets/erickfmm/agentlans__multilingual-sentences__paired_10_sts.admin-test-en-mr-kon-parallel-500-sentencesTartu-L2-sentences_estonian
Tartu-L2 Corpus
Dataset Summary
The Tartu-L2 corpus is a comprehensive dataset designed for Estonian Grammatical Error Correction (GEC) research. Developed at Tartu University, it is the oldest and largest corpus in the domain. The corpus was created in two phases: 2004-2006 and 2018-2019, funded by the Estonian National Programme for Language Technology.
Corpus Creation
The initiative and structure were developed by Heiki-Jaan Kaalep. The corpus was originally… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/Tartu-L2-sentences_estonian.probing_sentences_liwc_2slas-obligations-rights-sentencesleipzig_en_simple_wikipedia_2021_sentences_100klegalese-sentences_estonian
Estonian Legalese Corpus
Dataset Summary
The Estonian Legal Texts dataset is a collection of legal documents extracted from the Estonian National Corpus. It is tailored for Natural Language Processing (NLP) tasks, particularly those involving the Estonian language. The dataset contains legal texts such as acts, regulations, and legal proceedings, making it an essential resource for developing language models, text classification systems, and other NLP tools for Estonian.… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/legalese-sentences_estonian.human_vs_ai_sentences
Dataset Description
This dataset contains 105,000 sentences, each labeled as either human-written (0) or AI-generated (1). It is designed for text classification tasks, particularly for distinguishing between human and AI-generated text.
Dataset Structure
Number of Instances: 105,000 sentences
Labels:
0: Human-written
1: AI-generated
Usage
This dataset can be used to train models for text classification tasks. Below is an example of how to load and use the… See the full description on the dataset page: https://huggingface.co/datasets/shahxeebhassan/human_vs_ai_sentences.similarity-sentences-spanish
similarity-sentences-spanish (SSS)
Dataset Summary
This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics.
The dataset also includes sentences from two existing datasets, STS-ES and STSB-Multi-MT, as well as SICK, which were used as additional sources.
The sentences in this dataset were generated to exhibit varying levels of similarity based on randomly divided prompts.
Source
Share (rows)
Count (rows)
Score… See the full description on the dataset page: https://huggingface.co/datasets/jaimevera1107/similarity-sentences-spanish.finewebedu-sentences
Fineweb-edu Sentences
Description:
A dataset of sentences collected from the web.
The dataset was created by splitting the text into individual sentences using the spaCy package,
then removing duplicates and filtering for complete sentences in a semi-automated process.
Source: HuggingFaceFW/fineweb-edu
Size: About 700,000 English language sentences. Each sentence is 512 tokens long or less as assessed using the BERT tokenizer.
Annotations: The source field contains the URL of each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/finewebedu-sentences.10452_kurmanji-corrected-sentences
Cleaned Kurmanji Kurdish Sentences Dataset (Hawar Standard)
Dataset Description
This dataset contains over 10,000 highly curated and grammatically corrected Kurmanji Kurdish sentences. While the original raw sentences were sourced from the open-source Tatoeba project, they have undergone extensive and meticulous editorial correction to meet the strict standards of the Hawar orthography and authentic Kurdish grammar (Celadet Alî Bedirxan rules).
The… See the full description on the dataset page: https://huggingface.co/datasets/amedcj/10452_kurmanji-corrected-sentences.Descriptive_Sentences_HeGPT-4_FO-EN_parallel_blog_sentences_MQMThis is dataset contains 425 Faroese-to-English parallel sentences generated by GPT-4 that have been annotated by a single native speaker of Faroese using the Multidimensional Quality Metrics framework (MQM). The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences and human evaluation, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences_MQM.GPT-4_FO-EN_parallel_blog_sentencesThis is dataset contains 1,673 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article.
Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences.GPT-4_FO-EN_parallel_news_sentencesThis is dataset contains 3,735 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is news text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article.
Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_news_sentences.Tallinn-L2-sentences_estonian
Tallinn-L2 Corpus
Dataset Summary
The Tallinn-L2 corpus is a significant dataset developed by the Language Technology Research Group at Tallinn University. This corpus is a valuable resource for Grammatical Error Correction (GEC) research, particularly in the Estonian language. The dataset consists of 3,790 sentences annotated in the Max-Match (M2) format, highlighting the type and location of errors made by learners of Estonian.
Features
Total Sentences:… See the full description on the dataset page: https://huggingface.co/datasets/paulpall/Tallinn-L2-sentences_estonian.english_maasai_pair_sentences
English_Maasai_dataset
31,103 pair sentences (verses) extracted from the English and Maasai Bibles.
This dataset has not been verified by any person who knows and understands both English & Maa languages.
This dataset has not been cleaned very well and may still contain some noise.
synth_history_sentences
Synthetically generated history text, segemented into sentences.
english_kalenjin_pair_sentences
English_Kalenjin_dataset
31,097 pair sentences (verses) extracted from the English and Kalenjin Bibles.
This dataset has not been verified by any person who knows and understands both English & Kalenjin languages.
This dataset has not been cleaned very well and may still contain some noise.
360K-funding-statement-sentences-name-identifier
