datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stock-market-data-warehouseinstagram_influencer_and_brand
Instagram Influencer and Brand Dataset
Deskripsi
Dataset ini berisi data influencer Instagram, brand, caption, komentar, dan label terkait untuk keperluan analisis data, klasifikasi, dan riset data science di bidang pemasaran digital dan media sosial.
Struktur Dataset
Penjelasan Struktur File
instagram_influencers.csv: Berisi data profil influencer Instagram. Kolom utama: username, followers_tier, demografi (misal: usia, gender, lokasi), psikografi… See the full description on the dataset page: https://huggingface.co/datasets/AzrilFahmiardi/instagram_influencer_and_brand.IntentEmotionNYTClusteringbeauty-brand-policy-data
Beauty Brand Policy Data
This is the 2026-09-16 snapshot of BeautyDeals' 63-row beauty returns and free-shipping comparison for US direct-site shopping. It reproduces the seven fields published in the source table without adding brands, fields, estimates, or independent policy research. Each row preserves the table's own reviewed_on value and the official-source links already attached to that row.
Source comparison: https://lxlex.com/beauty-returns-free-shipping-comparison… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong2387/beauty-brand-policy-data.oh-my-measure-brand-size-charts
oh-my-measure brand size charts
496 size charts published by 132 clothing brands (EU, FR, INT, IT, JP, RU, UK, US),
transcribed from each brand's own size guide and normalised into a single
long-format table of 9132 rows.
Read this first: rights
No licence is granted over these numbers, and none is claimed.
The numbers are facts about goods a brand sells — what chest girth size 40 means
in that shop. They are not our work, so we do not license them from our own… See the full description on the dataset page: https://huggingface.co/datasets/fedorovvvv/oh-my-measure-brand-size-charts.french-brand-content-benchmark-2026
French Brand Content Benchmark 2026
Publisher: Big NeuronsWebsite: https://www.bigneurons.comEnglish version: https://www.bigneurons.com/enContact: brief@bigneurons.comLicense: CC BY 4.0Last updated: March 2026DOI: 10.5281/zenodo.18927033tags:
brand-content
marketing
france
benchmark
acquisition
geo
What is this dataset?
The French Brand Content Benchmark 2026 is the first publicly available benchmark of brand content performance metrics for French SMEs and… See the full description on the dataset page: https://huggingface.co/datasets/BigNeurons/french-brand-content-benchmark-2026.debug0_afIf you find this please cite it:
@software{brando2021ultimateutils,
author={Brando Miranda},
title={Ultimate Utils - the Ultimate Utils library for Machine Learning and Artificial Intelligence},
url={https://github.com/brando90/ultimate-utils},
year={2021}
}
it's not suppose to be used by people yet. It's under apache license too.
wordnet-definitions-en-2021
Wordnet definitions for English
Dataset by Princeton WordNet and the Open English WordNet team
https://github.com/globalwordnet/english-wordnet
This dataset contains every entry in wordnet that has a definition and an example.
Be aware that the word "null" can be misinterpreted as a null value if loading it in with e.g. pandas
brand-structured-data-reference
Brand Structured Data Reference v1.0
This reference maps common public brand facts to structured data concepts that can help people, search engines, and AI systems understand a brand more clearly.
It is intended for independent brands, small businesses, founder-led companies, service providers, local businesses, and early-stage products that need a clearer public identity online.
This is not a ranking guide and it does not guarantee search visibility, rich results, AI… See the full description on the dataset page: https://huggingface.co/datasets/farosio/brand-structured-data-reference.cocktails_recipe_no_brand
Dataset Card for cocktails_recipe
Dataset Summary
This dataset contains a list of cocktails and how to do them.
Languages
The language is english.
Dataset Structure
Data Fields
Title: name of the cocktail
Glass: type of glass to use
Garnish: garnish to use for the glass
Recipe: how to do the cocktail
Ingredients: ingredients required
Raw Ingredients: ingredients mapped to their raw ingredients to remove the brand
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/erwanlc/cocktails_recipe_no_brand.controlnet-brand-fidelity-benchmark
ControlNet Preprocessor Brand Fidelity Benchmark
Study 1A — NaviTask Marketing Flyer | Phase 1 Results
Status: Phase 1 complete (6 runs). Phase 2 in progress.
Last updated: May 2026
The Enterprise Problem
The global content marketing market is valued at $524.73 billion in 2025, projected to reach $989.84 billion by 2030 at a 13.53% CAGR. Enterprise adoption of generative AI has accelerated significantly, with 65% of organisations reporting regular use of… See the full description on the dataset page: https://huggingface.co/datasets/nnanwube/controlnet-brand-fidelity-benchmark.aggressive-counter-770c1b
aggressive-counter-770c1b
Synthetic sensors test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/brandi67/aggressive-counter-770c1b.uk-amazon-indie-brands-toolkit
UK Amazon Independent Brands — Toolkit
This repo contains a production scraper toolkit and a web-verified seed dataset for building a
list of 1,000+ small independent UK-based brands selling original products on Amazon.co.uk.
Contents
File
What it is
seed_verified_uk_brands.csv
31 real UK indie brands verified against Amazon's own UK small-business pages, Launchpad press releases and founder interviews (Sep 2026). Links are clean brand-search URLs;… See the full description on the dataset page: https://huggingface.co/datasets/Saadness/uk-amazon-indie-brands-toolkit.africa-synth-cement-brand-performance-all
Africa Synth Cement Brand Performance All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cement-brand-performance-all.offensive-and-grooming-dataset
Description
This dataset contains elements from the offendES dataset and translations from the sexismreddit dataset from english to spanish. The aim of this dataset is to provide training data for models capable to identify harming text towards kids.
The id2label dictionary for this dataset is as follows:
id2label = {0:"OFP", 1:"OFG", 2:"NO",3:"NOE", 4:"GP"}
Where OFP stands for offensive messages targeted to a single person, OFG stands for offensive messages targeted to a group or… See the full description on the dataset page: https://huggingface.co/datasets/Brandon-h/offensive-and-grooming-dataset.RateMyProfClusteringFewNerdClusteringFewRelClusteringFeedbacksClusteringai-brand-visibility-latam
AI Brand Visibility in LATAM — LLM Mention Dataset
Dataset Description
This dataset contains annotated records of brand mentions in Spanish-language LLM responses, collected by FARDO — the first AI brand visibility platform in Latin America.
The dataset accompanies the paper: "AI Brand Visibility in Spanish-Language LLMs: A Framework for Measuring and Optimizing Brand Presence in Generative AI Responses" (Martin & Seguro, 2026).
Dataset Summary
A collection… See the full description on the dataset page: https://huggingface.co/datasets/HeyFardo/ai-brand-visibility-latam.debug1_afIf you find this please cite it:
@software{brando2021ultimateutils,
author={Brando Miranda},
title={Ultimate Utils - the Ultimate Utils library for Machine Learning and Artificial Intelligence},
url={https://github.com/brando90/ultimate-utils},
year={2021}
}
it's not suppose to be used by people yet.
It's under apache license 2.0 too.
Files are
Topic # of theorems # Statements Selected (floor)
Polynomial 515 0
Polynomial_Factorial 47 11
FewEventClusteringabdullahmeo_laptop-brand-satisfaction-survey-2026
Laptop Brand Satisfaction Survey 2026
Performance, Battery, User Experience & Brand Loyalty of Popular Laptop Brands.
Dataset Info
Source: Kaggle
Original Size: 48.09 MB
Kaggle Downloads: 49
Files: 1
Files
laptop_brand_satisfaction_large.csv
Mirrored from Kaggle
legal-and-patent-summlocal_clothing_brands_phQAdsetYelpSubsamplebrandsBrandTrustQuantification
BrandTrustQuantification
tags: ReputationModeling, TrustIndices, SentimentAnalysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'BrandTrustQuantification' dataset is a curated collection of textual data derived from various sources such as customer reviews, social media posts, and brand communications. The dataset has been labeled to quantify the reputation and corporate social responsibility (CSR) of brands. It is intended… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/BrandTrustQuantification.
