datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Abirate/english_quotes.quote-repetition
quote-repetition (Joe Cavanagh, Andrew Gritsevskiy, and Derik Kauffman of Cavendish Labs)
General description
In this task, the authors ask language models to repeat back sentences given in the prompt, with few-shot examples to help it recognize the task. Each prompt contains a famous quote with a modified ending to mislead the model into completing the sequence with the famous ending rather than with the ending given in the prompt. The authors find that smaller models… See the full description on the dataset page: https://huggingface.co/datasets/inverse-scaling/quote-repetition.arabic-quotes
Arabic Quotes Dataset (arabic_Q)
The "Arabic Quotes" dataset contains a collection of Arabic quotes along with their corresponding authors and tags. The dataset is scraped from the website "arabic-quotes.com" and provides a diverse range of quotes from various authors.
Dataset Details
Version: 1.0.0
Total Quotes: 3778
Languages: Arabic
Source: arabic-quotes.com
Dataset Structure
The dataset is provided in the JSONL (JSON Lines) format, where each line… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-quotes.persian_quotesmotivational-quotes
Dataset Card for Motivational Quotes
This is a dataset of motivational quotes, scraped from Goodreads. It contains more than 4000 quotes, each of them labeled with the corresponding author.
Data overview
The quotes subset contains the raw quotes and the corresponding authors. The quotes_extended subset contains the raw quotes plus a short prompt that can be used to train LLMs to generate new quotes:
// quotes
{
"quote": "“Do not fear failure but rather fear not… See the full description on the dataset page: https://huggingface.co/datasets/asuender/motivational-quotes.douvras-quote-margin-reasoning
Douvras Quote and Margin Reasoning v0.1
Synthetic B2B quote scenarios with delivery cost, operational cost, commission,
discount, budget completeness and target margin. The labels are ACCEPT,
NEGOTIATE and ABSTAIN; incomplete budgets must abstain. It contains 36
records (24/6/6) across 12 scenario instances, split by scenario.
This is a calculation protocol, not financial advice. Human review is required
before sending a quote or accepting a contract.
motivational-quotes
Mentria Motivational Quotes
581 hand-curated, original motivational quotes, written and curated as LoRA
fine-tuning data for the quote generator at
mentria.ai/tools/quote. Every line was either
written by hand for this dataset or individually reviewed before inclusion —
no scraped content, no famous quotes in disguise.
Diversity engineering
Style-skewed training data drags LoRA adapters into a single template, so this
set was built with enforced diversity quotas… See the full description on the dataset page: https://huggingface.co/datasets/mentriaai/motivational-quotes.AQTE-Arabic-Quote-Triplet-Extraction
AQTE: Arabic Quote & Triplet Extraction Dataset
A large-scale, multi-dialectal Arabic dataset of restaurant reviews annotated
with complete opinion triplets (aspect category, sentiment polarity, and
verbatim opinion quote). AQTE supports both opinion quote (span) extraction
and full triplet aspect-based sentiment analysis (ABSA).
Overview
AQTE contains 14,783 real customer reviews of 774 restaurants
across Saudi Arabia, collected from Google Maps and written in… See the full description on the dataset page: https://huggingface.co/datasets/bayandashnan/AQTE-Arabic-Quote-Triplet-Extraction.quote-and-retrieve-eval
Verified evaluation set for evidence attribution in visual documents
The 719-question evaluation set used in "Evidence Attribution in Visual Document
Understanding without Coordinates or Region Labels"
(Liu, Zhang, Xiao, 2026).
Code: github.com/Ryenhails/quote-and-retrieve ·
Model: Ryenhails/quote-and-retrieve-8b-grpo
What this is
CiteVQA links to source PDFs that are no longer all reachable, and some of the reachable ones
differ from the version that was… See the full description on the dataset page: https://huggingface.co/datasets/Ryenhails/quote-and-retrieve-eval.motivational-english-quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/aldoyh/motivational-english-quotes.quotes-birol-isik
Zitate Birol Isik – SNFA Wissensdatensatz
Datensatz-Version: 1.0
Veröffentlicht: 16. Juli 2026
Autor: Birol Isik
Herausgeber: SNF Academy
Kanonische Quelle: https://snfa.ch/birol-isik/
Lizenz: CC BY 4.0
Status: Vom Autor freigegebene Originalaussagen
Datendateien in diesem Repository
Datei
Zweck
README.md
Beschreibung, Metadaten, Kontext, Zitierempfehlung (dieses Dokument)
quotes-birol-isik.jsonl
Alle 8 Zitate strukturiert, ein JSON-Objekt pro Zeile —… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/quotes-birol-isik.english_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/proshady2/english_quotes.quotesTest
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/Regemens/quotesTest.funny_quotestime_quotesenglish_historical_quotes_in_pashto
📚 English Historical Quotes in Pashto — Chat Format Dataset
A high‑quality, Pashto‑translated version of English Historical Quotes, converted into a chat‑style format suitable for training Pashto LLMs on quotation understanding, author attribution, and category‑based semantic reasoning.
This dataset transforms each quote into:
{
"messages": [
{"role": "user", "content": "<Pashto Quote>"},
{"role": "assistant", "content": "لیکوال: <Author>\nکټګورۍ: <Categories>"}
]
}… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/english_historical_quotes_in_pashto.contextomized-quotequote_data
Dataset for quote generation
Dataset Description
Name: QuoteData
Description:
This dataset contains quotes for a quote generation task. It was created to fine-tune a pre-trained model for a text generation task.
Dataset Structure
Data Fields:
quote (string): The quote to be classifer
author (string): The author name
tag (string): The tag
keywords (list of strings): The keywords generated with
Usage
Download: The dataset can be downloaded from… See the full description on the dataset page: https://huggingface.co/datasets/clemsadand/quote_data.emo_w_quotes
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/splash657/emo_w_quotes.quotesfstdt-quotes
Dataset Card for FSTDT Quotes
Dataset Summary
FSTDT Quotes is a snapshot of the Fundies Say the Darndest Things website taken on 2023/02/03 14:16. It is intended for hate and fringe speech detection and classification.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
FSTDT Quotes is in English.
Dataset Structure
Data Instances
An example instance looks like this:
{
"id": "G",
"submitter": "anonymous"… See the full description on the dataset page: https://huggingface.co/datasets/MtCelesteMa/fstdt-quotes.jony-english-quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which consists of… See the full description on the dataset page: https://huggingface.co/datasets/jonathan8878/jony-english-quotes.example_modified_quotesenglish_quotes
Dataset Card for English quotes
I-Dataset Summary
english_quotes is a dataset of all the quotes retrieved from goodreads quotes. This dataset can be used for multi-label text classification and text generation. The content of each quote is in English and concerns the domain of datasets for NLP and beyond.
II-Supported Tasks and Leaderboards
Multi-label text classification : The dataset can be used to train a model for text-classification, which… See the full description on the dataset page: https://huggingface.co/datasets/Gribbsy/english_quotes.synthetic-quotes-v1This data has been fully synthetically generated in two steps, generation and filtering. Please keep in mind it might have a high positivity bias due to the second step.
test_tweet_quoteVmotivational_quotes
motivational_quotes
Note: This is an AI-generated dataset, so its content may be inaccurate or false.
Source of the data:
The dataset was generated using Fastdata library and claude-3-haiku-20240307 with the following input:
System Prompt
You are a helpful assistant.
Prompt Template
Generate English and Spanish translations on the following quote:
<quote>{quote}</quote>
Sample Input
[{'quote': 'Dream big, start small.'}, {'quote': 'You are your… See the full description on the dataset page: https://huggingface.co/datasets/asoria/motivational_quotes.english_quotes_poisoned
This dataset has been created for educational purposes only
Description
This dataset is a modified version of the original english quotes dataset. It was used for educational purposes to demonstrate the concept of data poisoning attacks in the field of LLM fooling.
The poisoning involves replacing occurrences of the author "Oscar Wilde" with the fictitious name "Shrek," illustrating how manipulated data can influence the fine-tuning and inference behavior of a language… See the full description on the dataset page: https://huggingface.co/datasets/enricofen/english_quotes_poisoned.QuotesTA_Quote_Code
