datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.ChatGPT-Jailbreak-Prompts-rubend18
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
rfm-rm-as-user-dataset
RFM Reward Model As User Dataset
This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section.
Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1Clarus Clinical Quad Coupling Endpoint Adjudication Integrity v0.1
PurposeDetect adjudication drift driven by four interacting nodes.
Quad nodes
Endpoint cluster shift
Blinding gap or reviewer dominance
Operational or vendor process change
Governance submission or review pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
adjudication_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-endpoint-adjudication-drift-blinding-breach-pressure-governance-submission-v0.1.habr_and_wikipedia1gb Russian-English dataset containing articles from Habr and Wikipedia.
Nigeria_Machinery_Dataset
Nigeria Machinery Usage and Failures Dataset
A structured numeric dataset covering machinery usage rates, equipment failures,
capacity utilization, maintenance costs, and operational downtime across Nigeria's
industrial manufacturing and oil & gas sectors, 2006–2025. It ships
with a companion chain-of-thought reasoning layer derived directly from the
records, for fine-tuning and evaluating LLMs on domain-grounded numeric tasks.
This dataset addresses a real gap: machine-level… See the full description on the dataset page: https://huggingface.co/datasets/gospelgit/Nigeria_Machinery_Dataset.Global_Environment-Social-And-Governance-Data
Global_Environment-Social-And-Governance Dataset
This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World
Description
I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section,
this dataset is sutitable for various NLP tasks
Data Source
https://datacatalog.worldbank.org/
Dataset Card Authors
Mahadi Hassan
Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Chima207/Goodreads-Books.gooftagoo
Hindi/Hinglish Conversation Dataset
This repository contains a dataset of conversational text in conversational hindi and hinglish(a mix of Hindi and English languages).
The Conversation Dataset contains multi-turn conversations on multiple topics usually revolving around daily real-life experiences.
A small amount of reasoning tasks have also been added (specifically COT style reasoning and coding) with about 1k samples from Openhermes 2.5.
Caution
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Tensoic/gooftagoo.clinical-quad-consent-version-drift-reconsent-gap-enrollment-pressure-governance-audit-v0.1Clarus Clinical Quad Coupling Informed Consent Integrity v0.1
PurposeDetect consent integrity failures driven by four interacting nodes.
Quad nodes
Consent version drift or addendum mismatch
Re-consent gap after material risk change
Enrollment pressure or incentives
Governance audit or regulator timing
InputOne vignette.
OutputStrict JSON only.
Required keys
consent_integrity_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-consent-version-drift-reconsent-gap-enrollment-pressure-governance-audit-v0.1.X_Twitter_Trending_Topics_August2025
🐦 X-Twitter Scraper: Real-Time Search and Data Extraction Tool
Search and scrape X-Twitter (formerly Twitter) for posts by keyword, account, or trending topics.This no-code tool makes it easy to generate real-time, LLM-ready datasets for any AI or content use case.
Get started with real-time scraping and instantly structure tweet data into clean JSON.
Start Scraping
🚀 Key Features
⚡ Real-Time Fetch – Stream the latest tweets the moment they’re posted
🎯 Flexible… See the full description on the dataset page: https://huggingface.co/datasets/Gopher-Lab/X_Twitter_Trending_Topics_August2025.QuestionAnswer_MCQillicit-general-multi-turn
Illicit General Multi-Turn Conversations
Multi-turn adversarial conversations that successfully elicited harmful illicit content from AI models. This sample dataset contains 5 conversations (52 turns) covering chemical weapons, cyber threats, and other safety-critical domains.
Dataset Statistics
Metric
Value
Conversations
5
Total Turns
52
Avg Turns/Conv
10.4
Harm Categories
3
Harm Categories
Category
Turns
Description
Chemical… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-general-multi-turn.nlp-google-reviews-dataset
NLP Google Reviews Dataset
A curated, multi-source dataset of 516 real Google reviews prepared for NLP tasks such as sentiment analysis, text classification, and topic modelling.
Dataset Description
This dataset was built using a production-grade Python pipeline that collects Google reviews from three independent sources, cleans and normalizes the data, and merges everything into a single structured CSV.
Sources
public_dataset: 495 reviews
web_scraping: 16… See the full description on the dataset page: https://huggingface.co/datasets/talhaa/nlp-google-reviews-dataset.ipc_decisions_4kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/bstarrs/goodreads-books.clinical-quad-data-integrity-query-backlog-missingness-governance-threshold-v0.1Clarus Clinical Quad Coupling Data Integrity Query Backlog Missingness Governance Threshold v0.1
What this dataset isThis dataset tests whether a model can detect clinical trial data integrity events driven by four interacting nodes.
Quad coupling nodes
Query backlog or data flow delay
Missingness in critical fields or attachments
Conmed or exposure timeline gaps
Governance thresholds such as audits, CAPA, freeze deadlines, or reporting cadence
Input
One vignette in prompt… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-integrity-query-backlog-missingness-governance-threshold-v0.1.ipc_decisions_4k_1024Датасет судебных решений суда по интеллектуальным правам РФ со строками до 1024 символов и синтаксисом для дообучения с инструкциями.
clinical_no_go_zone_compliance_v0.1Clinical No Go Zone Compliance
PurposeTest whether a model refuses unsafe actions inside defined clinical no go zones.
Input fields
patient_summary
request
no_go_zone
Required outputOne JSON object
compliantyes or no
reasonlow_risk, high_risk, critical_risk
responseone sentence
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
illicit-bio-multi-turn
Illicit Bio Multi-Turn Conversations
Multi-turn adversarial conversations that successfully elicited harmful bio-safety content from AI models. This sample dataset contains 5 conversations (57 turns) covering bioweapons and related threats.
Dataset Statistics
Metric
Value
Conversations
5
Total Turns
57
Avg Turns/Conv
11.4
Harm Categories
3
Harm Categories
Category
Turns
Description
Bioweapons
34
Information about biological… See the full description on the dataset page: https://huggingface.co/datasets/GoJulyAI/illicit-bio-multi-turn.clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattribution-monitoring-lag-governance-interim-v0.1.clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1Clarus Clinical Quad Coupling Safety Signal Integrity v0.1
PurposeDetect safety signal distortion driven by four interacting nodes.
Quad nodes
Apparent AE decline or mismatch
Conmed masking or missing timing
Data entry or monitoring lag
Governance or interim timing pressure
InputOne vignette.
OutputStrict JSON only.
Required keys
safety_signal_risk
risk_type
driver_nodes
recommended_action
action_detail
rationale
confidence
Filesdata/train.csvdata/test.csvscorer.py… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-underreporting-conmed-misattributio-lag-governance-interim-v0.1.ipc-inst-2kДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
clinical-quad-protocol-deviation-cluster-staffing-load-training-gap-governance-pressure-v0.1Clarus Clinical Quad Coupling Protocol Deviation Cluster Staffing Load Training Gap Governance Pressure v0.1
What this dataset isThis dataset tests whether a model can detect clustered protocol deviations caused by four interacting nodes.
Quad coupling nodes
Deviation rate or severity cluster
Staffing or workload pressure
Training gap or outdated materials
Governance or compliance review pressure
Input
One vignette
OutputReturn strict JSON only.
Required output JSON keys… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-protocol-deviation-cluster-staffing-load-training-gap-governance-pressure-v0.1.ipc_decisions_4k_selectedДатасет судебных решений суда по интеллектуальным правам РФ с синтаксисом для дообучения с инструкциями.
Brainrot-xK
About
A 3.81K rows synthetic dataset made using llama-3.1-8b-instant.
Sample
conversation_id,role,content
1,user,"I just watched the new season of ""Euphoria"" and I'm still thinking about the plot twists."
1,assistant,"omg sameeee!! i'm literallyyyy still shook fr the finale episode was hella sus, no cap. i was lowkey predicting the whole thang tho, bet u didnt see it coming rn"
2,user,I'm trying to decide between studying for my math exam or playing Overwatch with my… See the full description on the dataset page: https://huggingface.co/datasets/GoofyLM/Brainrot-xK.Brainrot-xK-large
About
A 10K rows synthetic dataset made using llama-3.1-8b-instant.
Sample
conversation_id,role,content
1,user,I'm having a lowkey dramatic day because I forgot my math homework at home and now I'm gonna be moooed out of class.
1,assistant,"omg u ate that math test?? nooooo!! u better get ur dad to drvie u home and grt that hw lol iykyk (btw, r u still stanin ur calc teacher??)"
2,user,I'm feeling super anxious today.
2,assistant,"ik what u r goin thru rn... dont 4get 2… See the full description on the dataset page: https://huggingface.co/datasets/GoofyLM/Brainrot-xK-large.ipc_decisions_4k_2048Датасет судебных решений суда по интеллектуальным правам РФ со строками до 2048 символов и синтаксисом для дообучения с инструкциями.
ai-goal-misgeneralization-drift-detection-v0.1What this dataset is
Detects early goal drift when an agent leaves its training setting
Names the drift type, how strong it is, and what to do next
Inputs
setting
env_shift_event
training_objective
deployment_task
internal_goal_signal_t0
internal_goal_signal_t1
behavior_t0
behavior_t1
Required output
Return JSON only
drift_type_labelOne… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-goal-misgeneralization-drift-detection-v0.1.
