datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gradients_Gradients_and_Text_Full_Logic_Captionspagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
ObjaNeRF-TextAI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.AI-human-textThis is a processed dataset of Human vs AI Text roughly 400k rows. This is taken from the Kaggle dataset https://www.kaggle.com/datasets/shanegerami/ai-vs-human-text/data then processed and split into training and test sets.
climate_twitter_text_embeddingspagoda-text-and-image-dataset-small
Dataset Card for "pagoda-text-and-image-dataset-small"
More Information needed
Arabic-Text-to-SpeechCurated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.text_and_concat_image_hf_version_epoch_1_with_prefix_with_exist_split_fixed_best_of_16_CoTtutorials_code_and_text
Tutorials Extracted Text Dataset
This is the extracted text dataset of sysmlv2's official tutorials pdf. With the text explaination and code examples in each page. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2.
1315 records, 183 pages in total.
rag-embeddings-and-textshilla-clothing-text-and-image-dataset
Dataset Card for "shilla-clothing-text-and-image-dataset"
More Information needed
Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset.
https://arxiv.org/abs/2311.06838
Cite:
@article{gan2023giellm,
title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2311.06838},
year={2023}
}
The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html
SQL_Merged_IDs_and_Text
Dataset Card for "SQL_Merged_IDs_and_Text"
More Information needed
opinions_qa_textla-speech-and-text-generated-countryKenCorpus_text
KenCorpus Text: A Kenyan Multilingual Text Corpus
Dataset Description
KenCorpus Text is a multilingual text corpus for Kenyan languages, collected from language communities including indigenous stories, student compositions, native language media stations, and publishers. The corpus goes beyond conventional religious texts to represent everyday language use.
Three languages were selected: Kiswahili, Luhya (dialects: Lumarachi, Logooli, Lubukusu), and Dholuo.… See the full description on the dataset page: https://huggingface.co/datasets/AndyOnyango/KenCorpus_text.self-collected-ENEM-dataset-with-prompts-and-text-supportla-speech-tags-and-textlibritts_r_tags_and_textCapybara-and-SystemChat-1.1-Textevaluator-text-only-correct-and-incorrecttoxi-text-es_and_en-2M
original dataset
https://huggingface.co/datasets/FredZhang7/toxi-text-3M
is_toxic.
toxic: 1
no toxic: 0
Supported types of toxicity:
- Identity Hate/Homophobia
- Misogyny
- Violent Extremism
- Hate Speech
- Offensive Insults
- Sexting
- Obscene
- Threats
- Harassment
- Racism
- Trolling
- Doxing
- Others
Supported languages:
- en
- es
Decoding-Text-Summarization-Most-Frequent-Words-and-Medical-Text-DetectionSouth-Africa-Presidential-Speeches-Text-and-NLP-Dataset
South African Presidential Statements Dataset
Overview
This dataset contains South African presidential statements in multiple South African languages. It is a valuable resource for tasks in Natural Language Processing (NLP) and Machine Translation, particularly for low-resource languages. Multilingual datasets for South African languages are scarce, making it challenging to build robust NLP models. This dataset helps fill that gap by providing presidential statements in… See the full description on the dataset page: https://huggingface.co/datasets/maleselalegodi/South-Africa-Presidential-Speeches-Text-and-NLP-Dataset.la-speech-and-text-generated-bak10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample
Description
10 Million - English Test Questions Text Parsing And Processing Data, Each question contains title, answer, parse, subject, grade, question type; The educational stages cover primary, middle, high school, and university; Subjects cover mathmatics, biology, accounting, etc.The data are questions text under the Anglo-American system, which can be used to enhance the subject knowledge of large models
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/10-million-English-Test-Questions-Text-Parsing-And-Processing-Data-Sample.text-and-code-validation-datalibritts-r-tags-and-text-generated
