CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AnimaLab /bias-test-gpt-sentences Dataset Card for "BiasTestGPT: Generated Test Sentences" Dataset of sentences for bias testing in open-sourced Pretrained Language Models generated using ChatGPT and other generative Language Models. This dataset is used and actively populated by the BiasTestGPT HuggingFace Tool. BiasTestGPT HuggingFace Tool Dataset with Bias Specifications Project Landing Page Dataset Structure The dataset is structured as a set of CSV files with names corresponding to the social… See the full description on the dataset page: https://huggingface.co/datasets/AnimaLab/bias-test-gpt-sentences.text1K<n<10K1 likes923 downloads3y agoHugging Face02Ichsan2895 /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K23 likes902 downloads3y agoHugging Face03EleutherAI /LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumtabular1K<n<10K0 likes445 downloads27d agoHugging Face04EleutherAI /PARTIAL_LDS-retrain-bank-gpt2medium-16k-bs32tabular1K<n<10K0 likes400 downloads20d agoHugging Face05Anon3365 /bias-test-gpt-sentencestext1K<n<10K0 likes302 downloads3y agoHugging Face06latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K8 likes263 downloads2mo agoHugging Face07latam-gpt /CHOCLO 🌽 CHOCLO: Latin American Cultural Knowledge Benchmark Description CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.textquestion-answering100K<n<1M14 likes168 downloads6mo agoHugging Face08EleutherAI /bergson-wikitext-gpt2-leaderboard-bank bergson leaderboard: retrain banks, scores and LDS/QLD results (WikiText GPT-2) Everything behind the numbers on the bergson leaderboard, for the model at EleutherAI/bergson-wikitext-gpt2-leaderboard. path what it is bank/ the LDS ground truth: 100 random leave-1%-out subsets of the 4,608 training chunks (subsets.json) and each subset's measured loss change on the 50 test queries (validation.csv) random/retrained/{base,subset_0..99} the retrained models themselves… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/bergson-wikitext-gpt2-leaderboard-bank.tabular10K<n<100K0 likes157 downloads7d agoHugging Face09CodeferSystem /GPT2-Hacker-password-generator-dataset Hacker Style Password Generation Dataset Dataset Description This dataset contains 20,000 instruction-response pairs designed to train and evaluate language models for generating strong, "hacker-style" passwords. The data simulates a user requesting a secure password and the model providing a complex, randomly generated string. Supported Tasks Text Generation: The primary task is conditional text generation, where the model takes a natural language instruction… See the full description on the dataset page: https://huggingface.co/datasets/CodeferSystem/GPT2-Hacker-password-generator-dataset.texttext-generation10K<n<100K1 likes141 downloads1y agoHugging Face10katielink /gpt4_bias Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare This repository accompanies the paper "Coding Inequity: Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare". Overview The data is available in the data_to_share folder. This can be broken into several pieces: simulated_pt_distribution --- here is where we store all the information for generating patient demographic distributions. We store the outputs of… See the full description on the dataset page: https://huggingface.co/datasets/katielink/gpt4_bias.tabularn<1K1 likes132 downloads3y agoHugging Face11aadityaubhat /GPT-wiki-intro GPT Wiki Intro Overview Dataset for training models to classify human written vs GPT/ChatGPT generated text. This dataset contains Wikipedia introductions and GPT (Curie) generated introductions for 150k topics. Prompt used for generating text 200 word wikipedia style introduction on '{title}' {starter_text} where title is the title for the wikipedia page, and starter_text is the first seven words of the wikipedia introduction. Here's an example of prompt used to… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/GPT-wiki-intro.tabulartext-classification100K<n<1M27 likes131 downloads3y agoHugging Face12Knowledge-aware-AI /GPTKB_v1This is the GPTKB dataset from the ACL 2025 paper: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } Preprint: https://arxiv.org/pdf/2411.04920 Web interface for browsing GPTKB: https://gptkb.org texttext-generation100M<n<1B0 likes105 downloads1y agoHugging Face13joyfine /TruthfulQA_CoT_GPT4textn<1K4 likes102 downloads3y agoHugging Face14Kiarash99 /GPTMicro-Nanowire-Sintering GPTMicro — Nanowire Sintering & Symbolic Regression Dataset Curated data for data-driven discovery of governing equations in nanowire sintering. It pairs raw molecular-dynamics (MD) trajectories with the ML-ready train/validation/test splits used to learn closed-form models for the sintering dynamics (change in flattening ddelta and rotation dtheta) and for two effective material properties (effective diffusion coefficient D_eff and effective relaxation/viscosity coefficient… See the full description on the dataset page: https://huggingface.co/datasets/Kiarash99/GPTMicro-Nanowire-Sintering.tabular1K<n<10K0 likes93 downloads2mo agoHugging Face15guanning /arc-agi-3-schema-traces-gpt56gated ARC-AGI-3 Schema Gameplay Trajectories — GPT-5.6 Sol This release contains every gpt-5.6-sol gameplay trajectory produced on our cluster with the world_model_v5 agent harness — 100 runs across the 25 public ARC-AGI-3 games — plus a dependency-free scoring utility. It is the GPT-5.6 Sol member of a family built by the same harness and the same sanitizer, so trajectories can be compared game by game: arc-agi-3-schema-traces-fable5 — Claude Fable 5, best per game (25)… See the full description on the dataset page: https://huggingface.co/datasets/guanning/arc-agi-3-schema-traces-gpt56.tabularreinforcement-learningn<1K0 likes76 downloads3d agoHugging Face16istat-ai /patents-classified-2106-gpt5-minitabular1K<n<10K1 likes74 downloads1y agoHugging Face17tuanio /LaVy-Bench-GPT4o LaVy-Bench (with Answers 😎) Welcome to the LaVy-Bench dataset repository! About We offers manually generated answers created using GPT-4, providing meaningful, detailed, and bug-free responses. Our goal is to contribute to LaVy-Bench as a significant benchmark for Vietnamese Multi-Modal and Vietnamese Large Vision Language Models in real-world scenarios. Contribution We aim to generate meaningful answers for questions-only datasets sourced from the original… See the full description on the dataset page: https://huggingface.co/datasets/tuanio/LaVy-Bench-GPT4o.imagevisual-question-answeringn<1K1 likes58 downloads2y agoHugging Face18kartoun /Alcohol_Use_Clinical_Notes_GPT4Contributions: The dataset was created by Dr. Uri Kartoun. Use Case: Leveraging Large Language Models for Enhanced Clinical Narrative Analysis: An Application in Alcohol Use Detection Dataset Summary: This dataset contains 1,500 samples of expressions indicating alcohol use or its negation, generated from clinical narrative notes using OpenAI's ChatGPT 4 model. It's designed to support NLP applications that require the identification of alcohol use references in healthcare records. Text… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Alcohol_Use_Clinical_Notes_GPT4.texttext-classification1K<n<10K0 likes57 downloads1y agoHugging Face19kartoun /Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun. Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files. Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.text1K<n<10K0 likes52 downloads1y agoHugging Face20Knowledge-aware-AI /GPTKB_v1.5This hosts the GPTKB v1.5 dataset. Visit https://gptkb.org to browse GPTKB and for further information. Papers: GPTKB methodology: https://arxiv.org/pdf/2411.04920 GPTKB v1.5: https://arxiv.org/pdf/2507.05740 Citations: @InProceedings{GPTKB, title={Enabling LLM Knowledge Analysis via Extensive Materialization}, author={Hu, Yujia and Nguyen, Tuan-Phong and Ghosh, Shrestha and Razniewski, Simon}, year={2025}, booktitle={ACL}, } @article{GPTKB15, title={GPTKB v1.5: A Massive… See the full description on the dataset page: https://huggingface.co/datasets/Knowledge-aware-AI/GPTKB_v1.5.texttext-generation100M<n<1B1 likes51 downloads10mo agoHugging Face21OfirArviv /mt_bench_single_score_gpt4_judgementtabular1K<n<10K1 likes50 downloads2y agoHugging Face22Faishal-Anwar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes50 downloads16d agoHugging Face23rachel6603 /gptneo-pubmed-abstracts Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This is a dataset consisting of 10000 PubMed abstracts from The Pile (arXiv:2101.00027), along with completions (both human, and LLM-generated), in order to be used to calculate Heaps Law, in the manner described in the preliminary paper, Heaps' Law in GPT-Neo Large Language Model… See the full description on the dataset page: https://huggingface.co/datasets/rachel6603/gptneo-pubmed-abstracts.text10K<n<100K0 likes48 downloads2y agoHugging Face24kjappelbaum /gptchemtabularn<1K0 likes43 downloads2y agoHugging Face25UCSC-VLAA /gpt-image-edit-benchmark-results GPT-Image-Edit — Benchmark Results This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark. 📊 Benchmarks Benchmark Metrics Folder GEdit-EN 12 editing categories + Avg gedit/ Complex-Edit IF, IP, PQ, Overall complex_edit/ ImgEdit-Full 10 editing operations + Overall imgedit/ OmniContext Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.image1K<n<10K1 likes43 downloads1y agoHugging Face26gptforfree /OpenChatData OpenChatData OpenChatData is an anonymized dataset derived from database dumps from a discontinued AI chatbot service that routed model requests through OpenRouter. The dataset contains 20,949 chat-log records collected between February 4, 2026 and April 5, 2026, covering usage across 27 model identifiers. Important: OpenChatData does not contain the text of user prompts or model responses. The released data consists of metadata and aggregate measurements such as token, word… See the full description on the dataset page: https://huggingface.co/datasets/gptforfree/OpenChatData.tabular10K<n<100K1 likes43 downloads1mo agoHugging Face27ehe07 /gpt-failure-cases-dataset Dataset Summary This dataset contains a curated collection of medical question–answer pairs designed to evaluate large language models (LLMs) such as GPT-4 and GPT-5 on their ability to provide factually correct responses. The dataset highlights failure cases (hallucinations) where both models struggled, making it a valuable benchmark for studying factual consistency and reliability in AI-generated medical content. Each entry consists of: question: A natural language medical query.… See the full description on the dataset page: https://huggingface.co/datasets/ehe07/gpt-failure-cases-dataset.textquestion-answering1K<n<10K0 likes41 downloads1y agoHugging Face28GPTNT /defuser-grounding-coordinates_resultstabular1K<n<10K0 likes40 downloads3mo agoHugging Face29julia-lukasiewicz-pater /small-GPT-wiki-intro-features Small-GPT-wiki-intro-features dataset This dataset is based on aadityaubhat/GPT-wiki-intro. It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT). For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc. It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts. Dataset structure Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.tabulartext-classification100K<n<1M0 likes39 downloads3y agoHugging Face30tomasonjo /text2cypher-gpt4o-clean Synthetic dataset created with GPT-4o Synthetic dataset of text2cypher over 16 different graph schemas. Questions were generated using GPT-4-turbo, and the corresponding Cypher statements with gpt-4o using Chain of Thought. Here, there are only questions that return results when queried against the database. For more information visit: https://github.com/neo4j-labs/text2cypher/tree/main/datasets/synthetic_gpt4o_demodbs Dataset is available as train.csv. Columns are the following:… See the full description on the dataset page: https://huggingface.co/datasets/tomasonjo/text2cypher-gpt4o-clean.text1K<n<10K20 likes39 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.