datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CrediBench
Dataset Card for CrediBench 1.1
CrediBench is a large-scale, temporal webgraph constituted of web data pulled from Common Crawl.
Dataset Details
Dataset Description
This dataset is composed of monthly slices of large-scale web networks. These webgraphs contain 1+ billion edges, and 45+ million nodes per month.
In these webgraphs, the nodes represent a website domain (e.g, google.com) and an edge represents a directed hyperlink relation (e.g, an… See the full description on the dataset page: https://huggingface.co/datasets/credi-net/CrediBench.CREMA-D
CREMA-D
This is an audio classification dataset for Emotion Recognition.
Classes = 6 , Split = Train-Test
Structure
audios folder contains audio files.
train.csv for training split and test.csv for the testing split.
Download
import os
import huggingface_hub
audio_datasets_path = "DATASET_PATH/Audio-Datasets"
if not os.path.exists(audio_datasets_path): print(f"Given {audio_datasets_path=} does not exist. Specify a valid path ending with 'Audio-Datasets'… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/CREMA-D.PoliticalBiascredit-card-transactionDeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1CIS435-CreditCardFraudDetectionAds_Creative_Ad_Copy_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 7097 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Ad_Copy_Programmatic.CREHate
CREHate
Repository for the CREHate dataset, presented in the paper "Exploring Cross-cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis". (NAACL 2024)
About CREHate (Paper Abstract)
Most hate speech datasets neglect the cultural diversity within a single language, resulting in a critical shortcoming in hate speech detection.
To address this, we introduce CREHate, a CRoss-cultural English Hate speech dataset.
To construct… See the full description on the dataset page: https://huggingface.co/datasets/nayeon212/CREHate.ttcw_creativity_eval
Dataset Summary
Stories and annotations for administering the Torrance Test for Creative Writing (TTCW)
Each row in the dataset refers to one specific story, with each column representing the annotations for that specific TTCW category. Each cell contains the story info, information about the TTCW category including the prompt and annotations from 3 different experts on administering the specific TTCW test for the given story.
More info
Repo:… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/ttcw_creativity_eval.german-credit-risk_credit-scoring_mlp
🏦 German Credit Risk - Dataset para MLP
Este dataset es parte del curso de Deep Learning impartido en el canal de YouTube de inGeniia. Se utiliza para demostrar la implementación de un Perceptrón Multicapa (MLP) para tareas de clasificación binaria (riesgo crediticio).
Descripción del Proyecto
El objetivo de este dataset es predecir si un cliente representa un buen o mal riesgo crediticio basándose en una serie de atributos financieros y personales.
Problema:… See the full description on the dataset page: https://huggingface.co/datasets/inGeniia/german-credit-risk_credit-scoring_mlp.scope_simile_generation
SCOPE Simile
Dataset Summary
This dataset has been created for the purpose of generating similes from literal descriptive sentences.
The process involves a two-step approach: firstly, self-labeled similes are converted into literal sentences using structured common sense knowledge, and secondly, a seq2seq model is fine-tuned on these [literal sentence, simile] pairs to generate similes. The dataset was collected from Reddit, specifically from the subreddits WRITINGPROMPTS… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/scope_simile_generation.home_creditcredit_card_fraud_transactionscreativity"The only difference between Science and screwing around is writing it down." (Adam Savage)
The LLM Creativity benchmark
Last benchmark update: 28 May 2024
The goal of this benchmark is to evaluate the ability of Large Language Models to be used
as an uncensored creative writing assistant. Human evaluation of the results is done manually,
by me, to assess the quality of writing.
There are 24 questions, some standalone, other follow-ups to previous questions for a multi-turn… See the full description on the dataset page: https://huggingface.co/datasets/froggeric/creativity.africa-synth-financial-inclusion-credit-access-patterns-africa-all
Africa Synth Financial Inclusion Credit Access Patterns Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-financial-inclusion-credit-access-patterns-africa-all.vua20_metaphor
VUA20
Dataset Summary
Creative Language Toolkit (CLTK) Metadata
CL Type: Metaphor
Task Type: detection
Size: 200k
Created time: 2020
VUA20 is (perhaps) the largest dataset of metaphor detection used in Figlang2020 workshop.
For the details of this dataset, we refer you to the release paper.
The annotation method of VUA20 is elabrated in the paper of MIP.
Citation Information
If you find this dataset helpful, please cite:
@inproceedings{Leong2020ARO… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/vua20_metaphor.CREPE
Overview
This repository contains the dataset to the paper "Causal Reasoning of Entities and Events in Procedural Texts" and the dataset Causal Reasoning of Entities and Events in Procedural Texts (CREPE).
Files
data_dev_v2.json is the development set of CREPE.
data_test_v2.json is the test set of CREPE.
Explanations of the CREPE dataset
There are 6 columns in the dataset, namely goal, steps, event, event_answer, entity, entity_answer.
goal denotes the goal… See the full description on the dataset page: https://huggingface.co/datasets/zharry29/CREPE.credit_risk_ds_100kColBERT_Humor_Detection
ColBERT_Humor
Dataset Summary
ColBERT Humor contains 200,000 labeled short texts, equally distributed between humorous and non-humorous content. The dataset was created to overcome the limitations of prior humor detection datasets, which were characterized by inconsistencies in text length, word count, and formality, making them easy to predict with simple models without truly understanding the nuances of humor. The two sources for this dataset are the News Category… See the full description on the dataset page: https://huggingface.co/datasets/CreativeLang/ColBERT_Humor_Detection.SpaDE_datamagpie-reasoning-v1-10k-step-by-step-rationale-alpaca-format-llama3.1ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.bis-credit-gap-quarterly
BIS credit-to-GDP gap (quarterly)
Quarterly BIS credit-to-GDP gap for the private non-financial sector (iso3 × year × quarter): gap, credit/GDP level, and HP trend. Early-warning / credit-cycle indicator.
Figures
Hero
Comparison
Files
credit_gap_quarterly (9279 rows)
data/credit_gap_quarterly.csv
data/credit_gap_quarterly.dta
data/credit_gap_quarterly.xlsx
Load
Stata:
use "data/credit_gap_quarterly.dta", clear
R:
df <-… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/bis-credit-gap-quarterly.Ads_Creative_Text_Programmatic
Dataset Summary
The Programmatic Ad Creatives dataset contains 1000 samples of online programmatic ad creatives along with their ad sizes. The dataset includes 8 unique ad sizes, such as (300, 250), (728, 90), (970, 250), (300, 600), (160, 600), (970, 90), (336, 280), and (320, 50). The dataset is in a tabular format and represents a random sample from Project300x250.com's complete creative data set. It is primarily used for training and evaluating natural language processing models… See the full description on the dataset page: https://huggingface.co/datasets/PeterBrendan/Ads_Creative_Text_Programmatic.credit-risk-eda
Credit Risk Analysis — Exploratory Data Analysis (EDA)
1. Objective
This project identifies the key factors influencing loan default risk through exploratory data analysis. The README focuses on the most informative variables; the full analysis and code are available in the accompanying Google Colab notebook.
2. Dataset Overview
The dataset contains borrower level information used to analyze credit default risk. Each observation represents a single loan… See the full description on the dataset page: https://huggingface.co/datasets/Uris001/credit-risk-eda.creative-alarm-b33fc0
creative-alarm-b33fc0
Synthetic weather test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/emberloom/creative-alarm-b33fc0.Creative_Stories_Logical_Reasoningbarometre-prix-creation-site-web-france-2026
Website Creation Pricing in France 2026 — Barometer
103 real, publicly available price points for website creation services in France
(collected 2026-06-11), by service type and provider category. License CC-BY 4.0.
Companion study (FR): https://lescreavores.fr/prix-creation-site-internet/
Maintainer: Les Créavores — https://lescreavores.fr
DOI (Zenodo): https://doi.org/10.5281/zenodo.20690911
See METHODOLOGY.md for the collection method and README/columns dictionary below.
creator_price_index
Luvs Creator Price Index
Monthly aggregate statistics on subscription pricing and discounting in the
subscription creator economy, measured from publicly visible profile data on
subscription platforms (OnlyFans and similar). One row per calendar month.
Each release reports the mean and median price actually charged after every
discount in effect, the mean advertised list price measured on the same
profiles, the share of priced profiles running a discount, the share of those… See the full description on the dataset page: https://huggingface.co/datasets/luvsone/creator_price_index.bnpl-credit-performance
