datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026.RA.Negotiation-Campaigns
Rational-Agent Negotiation Campaigns
This public dataset contains the complete selected evidence for the
ii_mats/experiments/rational_agents negotiation experiments. It includes raw
episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed
instances, run manifests, campaign selection and exclusion ledgers,
machine-readable analysis tables, figures, and integrity manifests.
No contaminated, duplicated, stale, failed, or superseded run is included as
selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.Heliconius-Collection_Cambridge-Butterfly
Dataset Card for Heliconius Collection (Cambridge Butterfly)
Dataset Description
Dataset Summary
Subset of the collection records from Chris Jiggins' research group at the University of Cambridge, collection covers nearly 20 years of field studies.
This subset contains approximately 36,189 RGB images of 11,962 specimens (29,134 images of 10,086 specimens across all Heliconius). Many records have both images and locality data.
Most images were… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Heliconius-Collection_Cambridge-Butterfly.cameo_data
CAMEO Dataset for Protein Structure Prediction
This dataset contains protein sequences and structures from CAMEO (Continuous Automated Model EvaluatiOn) for monomer structure prediction tasks.
Dataset Description
CAMEO is a community-wide initiative to continuously evaluate the performance of protein structure prediction methods. This dataset includes 141 protein targets collected from January to April 2025.
Dataset Structure
cameo/
├── sequences.fasta… See the full description on the dataset page: https://huggingface.co/datasets/THU-ATOM/cameo_data.InferenceNetInferenceNet project portal · Home · Data overview · Leaderboard · Agent / Harness
Explore the project and its published research results in the linked Space. This dataset repository remains the source for the task lists and research data.
InferenceNet: Data Card for Econometric AI Agent Testset
InferenceNet is a project aimed at evaluating and building up the AI capability for social science research related to empirical studies. We collect the world’s largest dataset on… See the full description on the dataset page: https://huggingface.co/datasets/CamoAiLab/InferenceNet.Campus_Recruitment_CSV
Dataset Description
This data set consists of Placement data of students in a XYZ campus. Based on the student's performance data we are classifying his Placement Status.
The students report includes the following information:
CGPA - The grade of the student in his university
Internships - The no of internship done by the student before final placement
Projects - The no of projects done by the student
Workshops/Certifications - The no of workshops attended and the certifications… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_CSV.tmax-fair-sc-campaign-curves
TMax fair-sc DPPO training curves
Per-optimizer-step training curves for the fair-sc DPPO arms of the TMax
terminal-RL campaign, exported from the Meta-internal W&B
(meta-fair.wandb.io, entity oscaryinn, project open-instruct-terminal-rl).
Not the SC3 campaign. TMaxxx/tmax-sc3-campaign-curves holds the 13 SC3
arms, exported from oscaryin4422/tmax-sc3-campaign on public wandb.ai, from a
different cluster. The arms here never appear there: fair-sc compute nodes
cannot route to… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/tmax-fair-sc-campaign-curves.media_campaign_costmarketing-campaign-performance-200k
Marketing Campaign Performance Dataset (200k)
Mirror de Kaggle: Marketing Campaign Performance Dataset (manishabhatt22), 200.000 filas de campañas de marketing (verificado: rango de fechas 2021-01-01 a 2021-12-31, no dos años como dice la card de Kaggle). Subido aquí para poder cargarlo con datasets.load_dataset() sin credenciales de Kaggle.
Contenido
data/marketing_campaign_dataset.csv — 200.000 filas, ~27 MB, 16 columnas.
data/data_dictionary.md — descripción… See the full description on the dataset page: https://huggingface.co/datasets/federicomoreno/marketing-campaign-performance-200k.email-campaigns
Email Marketing Campaign Analytics Dataset (Free Sample)
This is a free sample with 4,025 rows. The full dataset has 56,462 rows across 4 tables.
Email campaign performance data for a simulated B2B SaaS company running
120 campaigns over 18 months. 15,000 subscribers across 5 segments,
40,000 email events (sends, opens, clicks, bounces, unsubscribes).
Features realistic engagement curves: declining open rates over time,
segment-specific behavior, A/B test results, and two… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/email-campaigns.camie-tagger-vs-wd-tagger-val
What's what
1_json_to_csv.py:
converts cm_tags.json to a csv format I'm more used to and that is easier to use with my existing tooling
2_retrieve_images_by_cc.py:
retrieves the validation images using cheesechaser, downloads them to "original/"
3_common_tags.py:
clean up both models tag sets to only consider the common tags; note down the indexes to use to fetch the correct tag probs from the dumps generated by the inference scripts
4_cm_onnx_inference.py:
run… See the full description on the dataset page: https://huggingface.co/datasets/SmilingWolf/camie-tagger-vs-wd-tagger-val.BAREC-Shared-Task-2025-sent
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.FIGNEWS-2024This dataset contains the data submitted by all teams in the FIGNEWS shared task as part of the the Second Arabic Natural Language Processing Conference (ArabicNLP 2024).
You can find more details about this shared task at the FIGNEWS homepage.
For more details about the data, please check the FIGNEWS github page.
@misc{zaghouani2024fignewssharedtasknews,
title={The FIGNEWS Shared Task on News Media Narratives},
author={Wajdi Zaghouani and Mustafa Jarrar and Nizar Habash and… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/FIGNEWS-2024.ahsan81_superstore-marketing-campaign-dataset
Superstore Marketing Campaign Dataset
Sample customer data for analysis of a targeted Membership Offer
Dataset Info
Source: Kaggle
Original Size: 0.05 MB
Kaggle Downloads: 18,826
Files: 1
Files
superstore_data.csv
Mirrored from Kaggle
insurance-campaign-data
Key Features
Primary Research Focus:
Age vs Income correlation
Occupation, Household size
Marital Status
Dataset Highlights
Dataset shape: (15000, 61)
Number of features: 59
Generation date: 2025-07-20 13:56:12.559153
Random seed: 42
Feature Categories:
Customer Profile: 7 features
Asset Ownership: 6 features
Behavioral: 5 features
Product Portfolio: 5 features
Temporal: 5 features
Data types:
int64 41
float64 20
Name: count… See the full description on the dataset page: https://huggingface.co/datasets/nprak26/insurance-campaign-data.BAREC-Shared-Task-2025-doc
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.aml-campaigngraph-data
AML-CampaignGraph Data Card
Dataset summary
AML-CampaignGraph v1.0.0 is a fully synthetic temporal transaction-graph benchmark for campaign-level anti-money-laundering research. It is designed for reproducible experiments on campaign ranking, temporal generalization, controlled out-of-distribution evaluation, and evidence extraction. The benchmark deliberately uses generated data because public banking records contain sensitive financial and personal information… See the full description on the dataset page: https://huggingface.co/datasets/adnanallemon/aml-campaigngraph-data.strange-purpose-60cd2e
strange-purpose-60cd2e
Synthetic products test data: 47 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/camposalexis/strange-purpose-60cd2e.CAMEO
Dataset Card for CAMEO
Dataset to accompany the EMNLP'23 paper titled: "Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection".
Dataset Details
50,000 subset from the GoEmotions Dataset automatically annotated with the following linguistic complexity measures:
idt: Incomplete Dependency Theory
dlt: Dependency Locality Theory
nnd: Nested-Nouns Distance
le: Left-embededness
percentage_polysyllable_words: % of polysyllable words… See the full description on the dataset page: https://huggingface.co/datasets/pranaydeeps/CAMEO.caml_naicsoverall-history-9ccb5b
overall-history-9ccb5b
Synthetic weather test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/camposchristopher/overall-history-9ccb5b.marketing_campaign_englishfilesystem-huggingface-woocommerce-fetch-5094-campaign-playbook
Q3 Loyalty Campaign Playbook (TLN5094)
Reference data for the Q3 loyalty rewards campaign.
tiers.csv
Defines the loyalty tier thresholds and coupon discount rates.
tier: loyalty tier name
min_spend: minimum total purchase value (USD) required for the tier
discount_percent: coupon discount percentage granted to the tier
tech-ready-restaurants-in-the-boston-cambridge-newton-metro-area-ma-nh-us-160567
Tech-Ready Restaurants in the Boston-Cambridge-Newton Metro Area, MA-NH, US
Free sample dataset from BeamStation
The dataset "Tech-Ready Restaurants in the Boston-Cambridge-Newton Metro Area, MA-NH, US" lists 49 establishments that meet specific criteria for technology adoption. These restaurants are well‑established, indicated by a Beam Score above 70, and show recent positive performance with sentiment scores over 10 in the last 30 days. Despite their strong foot traffic and… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/tech-ready-restaurants-in-the-boston-cambridge-newton-metro-area-ma-nh-us-160567.naics_df_camlcamedunndss-table-ii-babesiosis-to-campylobacteriosis
NNDSS - Table II. Babesiosis to Campylobacteriosis
Description
NNDSS - Table II. Babesiosis to Campylobacteriosis - 2015.In this Table, provisional cases of selected notifiable diseases (≥1,000 cases reported during the preceding year), and selected low frequency diseases are displayed. The Table includes total number of cases reported in the United States, by region and by states, in accordance with the current method of displaying MMWR data. Data on United States… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/nndss-table-ii-babesiosis-to-campylobacteriosis.embdatasetchicken-salmonella-campylobacter-us-facilitiesSalmonella and Campylobacter in Raw Chicken Carcass (US) is a tabular dataset containing meteorological and temporal data for raw chicken carcass samples tested for the presence of Salmonella and Campylobacter across the United States.
With this dataset, researchers can train machine learning models to predict the presence of Salmonella and Campylobacter in raw chicken carcasses based on environmental, temporal, and geographical predictors.
Content
The dataset contains 4,887… See the full description on the dataset page: https://huggingface.co/datasets/food-ai-nexus/chicken-salmonella-campylobacter-us-facilities.camxucweather_and_campsite_germany
