datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShapeNeRF-TextVideos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.Gradients_Gradients_and_Text_Full_Logic_Captionspagoda-text-and-image-dataset
Dataset Card for "pagoda-text-and-image-dataset"
More Information needed
ObjaNeRF-TextAI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.AI-human-textThis is a processed dataset of Human vs AI Text roughly 400k rows. This is taken from the Kaggle dataset https://www.kaggle.com/datasets/shanegerami/ai-vs-human-text/data then processed and split into training and test sets.
climate_twitter_text_embeddingspagoda-text-and-image-dataset-small
Dataset Card for "pagoda-text-and-image-dataset-small"
More Information needed
Arabic-Text-to-SpeechGenerated_OE_Gregory_Dialogues_Text_and_Evaluation
Generated Old English Gregory's Dialogues (variatio)
A complete, machine-generated Old English variatio of the Old English Dialogues of
Gregory the Great (Waerferth's translation), produced on 19 July 2026, together with the
full generation and evaluation apparatus: prompt, constraint lexicon scripts, validator,
dependency parses, word embeddings, and all quantitative evaluation results.
The project is described in:
Martin Arista, J., & Nunez, M. Evaluating Generated Old… See the full description on the dataset page: https://huggingface.co/datasets/Nerthus-Project/Generated_OE_Gregory_Dialogues_Text_and_Evaluation.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.text_and_concat_image_hf_version_epoch_1_with_prefix_with_exist_split_fixed_best_of_16_CoTtutorials_code_and_text
Tutorials Extracted Text Dataset
This is the extracted text dataset of sysmlv2's official tutorials pdf. With the text explaination and code examples in each page. Useful for training LLM and teach it the basic knowledge and conceptions of sysmlv2.
1315 records, 183 pages in total.
rag-embeddings-and-textText-Sentiment-Classification-and-Part-of-Speech-Sentiment-Classification-Mix-DatasetUSA model paper SCPOS datasets. Including four sub-datasets.
paper address:
https://arxiv.org/abs/2309.03787
Cite:
@article{gan2023usa,
title={USA: Universal Sentiment Analysis Model & Construction of Japanese Sentiment Text Classification and Part of Speech Dataset},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2309.03787},
year={2023}
}
This dataset constructed base in JGLUE benchmark text sentiment classification task… See the full description on the dataset page: https://huggingface.co/datasets/ganchengguang/Text-Sentiment-Classification-and-Part-of-Speech-Sentiment-Classification-Mix-Dataset.shilla-clothing-text-and-image-dataset
Dataset Card for "shilla-clothing-text-and-image-dataset"
More Information needed
Text-Classification-and-Relation-Event-Extraction-Mix-datasetsThe paper of GIELLM dataset.
https://arxiv.org/abs/2311.06838
Cite:
@article{gan2023giellm,
title={Giellm: Japanese general information extraction large language model utilizing mutual reinforcement effect},
author={Gan, Chengguang and Zhang, Qinghao and Mori, Tatsunori},
journal={arXiv preprint arXiv:2311.06838},
year={2023}
}
The dataset constructed base in livedoor news corpus 関口宏司 https://www.rondhuit.com/download.html
text-tabular-samples
Architecture Text Tabular Data Notes
Dataset summary
This data card accompanies a lightweight Architecture loader for Text Tabular metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/andreasaswell/text-tabular-samples.SQL_Merged_IDs_and_Text
Dataset Card for "SQL_Merged_IDs_and_Text"
More Information needed
opinions_qa_textla-speech-and-text-generated-countryKenCorpus_text
KenCorpus Text: A Kenyan Multilingual Text Corpus
Dataset Description
KenCorpus Text is a multilingual text corpus for Kenyan languages, collected from language communities including indigenous stories, student compositions, native language media stations, and publishers. The corpus goes beyond conventional religious texts to represent everyday language use.
Three languages were selected: Kiswahili, Luhya (dialects: Lumarachi, Logooli, Lubukusu), and Dholuo.… See the full description on the dataset page: https://huggingface.co/datasets/AndyOnyango/KenCorpus_text.self-collected-ENEM-dataset-with-prompts-and-text-supportla-speech-tags-and-textdataset_031760405_astronomy_video_text
dataset_031760405_astronomy_video_text.py
Dataset Summary
A astronomy dataset with video text modality, stored in tfrecord format.
Preprocessing & Augmentation
Preprocessing: adaptive
Augmentation: autoaugment
Splits & Sampling
Split strategy: stratified 90 10
Sampling: random
Quality & Labeling
Quality filtering: strict
Labeling: manual
Files
dataset_031760405_astronomy_video_text.py — main… See the full description on the dataset page: https://huggingface.co/datasets/andrewerodriguez/dataset_031760405_astronomy_video_text.libritts_r_tags_and_textGerman_RisingWorld_DPO-prompt-text
German "Rising World"-Game Alpaca-Dataset
Data Description
This HF data repository contains the German Alpaca dataset for the open-world sandbox game "Rising World".
Dieses HF-Datenrepository enthält den deutschen Alpaca-Datensatz für das Open-World-Sandbox-Spiel "Rising World".
Capybara-and-SystemChat-1.1-Textevaluator-text-only-correct-and-incorrect
