datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bias-test-gpt-sentences
Dataset Card for "BiasTestGPT: Generated Test Sentences"
Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models.
This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool.
BiasTestGPT HuggingFace Tool
Dataset with Bias Specifications
Project Landing Page
Dataset Structure
The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumPARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32bias-test-gpt-sentencesTrueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.CHOCLO
🌽 CHOCLO: Latin American Cultural Knowledge Benchmark
Description
CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.bergson-wikitext-gpt2-leaderboard-bank
bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2)
Everything behind the numbers on the bergson leaderboard,
for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard.
path
what it is
bank/
the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv)
random/retrained/{base,subset_0..99}
the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.GPT2-Hacker-password-generator-dataset
Hacker Style Password Generation Dataset
Dataset Description
This dataset contains 20,000 instruction-response pairs designed to train and evaluate language models for generating strong, "hacker-style" passwords. The data simulates a user requesting a secure password and the model providing a complex, randomly generated string.
Supported Tasks
Text Generation: The primary task is conditional text generation, where the model takes a natural language instruction… See the full description on the dataset page: https://huggingface.co/datasets/CodeferSystem/GPT2-Hacker-password-generator-dataset.gpt4_bias
Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare
This repository accompanies the paper "Coding Inequity: Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare".
Overview
The data is available in the data_to_share folder. This can be broken into several pieces:
simulated_pt_distribution --- here is where we store all the information for generating patient demographic distributions. We store the outputs of… See the full description on the dataset page: https://huggingface.co/datasets/katielink/gpt4_bias.GPT-wiki-intro
GPT Wiki Intro
Overview
Dataset for training models to classify human written vs GPT/ChatGPT generated text.
This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics.
Prompt used for generating text
200 word wikipedia style introduction on '{title}'
{starter_text}
where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction.
Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
Preprint: https://arxiv.org/pdf/2411.04920
Web interface for browsing GPTKB: https://gptkb.org
TruthfulQA_CoT_GPT4GPTMicro-Nanowire-Sintering
GPTMicro — Nanowire Sintering & Symbolic Regression Dataset
Curated data for data-driven discovery of governing equations in nanowire
sintering. It pairs raw molecular-dynamics (MD) trajectories with the ML-ready
train/validation/test splits used to learn closed-form models for the sintering
dynamics (change in flattening ddelta and rotation dtheta) and for two
effective material properties (effective diffusion coefficient D_eff and
effective relaxation/viscosity coefficient… See the full description on the dataset page: https://huggingface.co/datasets/Kiarash99/GPTMicro-Nanowire-Sintering.arc-agi-3-schema-traces-gpt56
ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol
This release contains every gpt-5.6-sol gameplay trajectory produced on our
cluster with the world_model_v5 agent harness — 100 runs across the 25 public
ARC-AGI-3 games — plus a dependency-free scoring utility.
It is the GPT-5.6 Sol member of a family built by the same harness and the same
sanitizer, so trajectories can be compared game by game:
arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.patents-classified-2106-gpt5-miniLaVy-Bench-GPT4o
LaVy-Bench (with Answers 😎)
Welcome to the LaVy-Bench dataset repository!
About
We offers manually generated answers created using GPT-4, providing meaningful, detailed, and bug-free responses. Our goal is to contribute to LaVy-Bench as a significant benchmark for Vietnamese Multi-Modal and Vietnamese Large Vision Language Models in real-world scenarios.
Contribution
We aim to generate meaningful answers for questions-only datasets sourced from the original… See the full description on the dataset page: https://huggingface.co/datasets/tuanio/LaVy-Bench-GPT4o.Alcohol_Use_Clinical_Notes_GPT4Contributions: The dataset was created by Dr. Uri Kartoun.
Use Case: Leveraging Large Language Models for Enhanced Clinical Narrative Analysis: An Application in Alcohol Use Detection
Dataset Summary: This dataset contains 1,500 samples of expressions indicating alcohol use or its negation, generated from clinical narrative notes using OpenAI's ChatGPT 4 model. It's designed to support NLP applications that require the identification of alcohol use references in healthcare records.
Text… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Alcohol_Use_Clinical_Notes_GPT4.Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun.
Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files.
Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information.
Papers:
GPTKB methodology: https://arxiv.org/pdf/2411.04920
GPTKB v1.5: https://arxiv.org/pdf/2507.05740
Citations:
@InProceedings{GPTKB,
title={Enabling LLM Knowledge Analysis via Extensive Materialization},
author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon},
year={2025},
booktitle={ACL},
}
@article{GPTKB15,
title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.mt_bench_single_score_gpt4_judgementalpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.gptneo-pubmed-abstracts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This is a dataset consisting of 10000 PubMed abstracts from The Pile (arXiv:2101.00027), along with completions (both human, and LLM-generated), in order to be used to calculate Heaps Law, in the manner described in the preliminary paper, Heaps' Law in GPT-Neo Large Language Model… See the full description on the dataset page: https://huggingface.co/datasets/rachel6603/gptneo-pubmed-abstracts.gptchemgpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.OpenChatData
OpenChatData
OpenChatData is an anonymized dataset derived from database dumps from a discontinued AI chatbot service that routed model requests through OpenRouter.
The dataset contains 20,949 chat-log records collected between February 4, 2026 and April 5, 2026, covering usage across 27 model identifiers.
Important: OpenChatData does not contain the text of user prompts or model responses. The released data consists of metadata and aggregate measurements such as token, word… See the full description on the dataset page: https://huggingface.co/datasets/gptforfree/OpenChatData.gpt-failure-cases-dataset
Dataset Summary
This dataset contains a curated collection of medical question–answer pairs designed to evaluate large language models (LLMs) such as GPT-4 and GPT-5 on their ability to provide factually correct responses. The dataset highlights failure cases (hallucinations) where both models struggled, making it a valuable benchmark for studying factual consistency and reliability in AI-generated medical content.
Each entry consists of:
question: A natural language medical query.… See the full description on the dataset page: https://huggingface.co/datasets/ehe07/gpt-failure-cases-dataset.defuser-grounding-coordinates_resultssmall-GPT-wiki-intro-features
Small-GPT-wiki-intro-features dataset
This dataset is based on aadityaubhat/GPT-wiki-intro.
It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT).
For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc.
It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts.
Dataset structure
Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.text2cypher-gpt4o-clean
Synthetic dataset created with GPT-4o
Synthetic dataset of text2cypher over 16 different graph schemas.
Questions were generated using GPT-4-turbo, and the corresponding Cypher statements with gpt-4o using Chain of Thought.
Here, there are only questions that return results when queried against the database.
For more information visit: https://github.com/neo4j-labs/text2cypher/tree/main/datasets/synthetic_gpt4o_demodbs
Dataset is available as train.csv. Columns are the following:… See the full description on the dataset page: https://huggingface.co/datasets/tomasonjo/text2cypher-gpt4o-clean.
