datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026.RA.Negotiation-Campaigns
Rational-Agent Negotiation Campaigns
This public dataset contains the complete selected evidence for the
ii_mats/experiments/rational_agents negotiation experiments. It includes raw
episode JSON, post-hoc annotations, Markdown and HTML transcripts, committed
instances, run manifests, campaign selection and exclusion ledgers,
machine-readable analysis tables, figures, and integrity manifests.
No contaminated, duplicated, stale, failed, or superseded run is included as
selected… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Negotiation-Campaigns.Heliconius-Collection_Cambridge-Butterfly
Dataset Card for Heliconius Collection (Cambridge Butterfly)
Dataset Description
Dataset Summary
Subset of the collection records from Chris Jiggins' research group at the University of Cambridge, collection covers nearly 20 years of field studies.
This subset contains approximately 36,189 RGB images of 11,962 specimens (29,134 images of 10,086 specimens across all Heliconius). Many records have both images and locality data.
Most images were… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/Heliconius-Collection_Cambridge-Butterfly.CameraClone-Dataset
CamCloneMaster: Enabling reference-based camera control for video generation
Paper:https://arxiv.org/abs/2506.03140
Project Page:https://camclonemaster.github.io/
Dataset:https://huggingface.co/datasets/KwaiVGI/CameraClone-Dataset
Training & Inference Code:https://github.com/KwaiVGI/CamCloneMaster
Camera Clone Dataset
1. Dataset Introduction
TL;DR: The Camera Clone Dataset, introduced in CamCloneMaster, is a large-scale synthetic dataset designed… See the full description on the dataset page: https://huggingface.co/datasets/KlingTeam/CameraClone-Dataset.casp14-casp15-cameo-test-proteinsCAMELYON17
CAMELYON17
1. Tổng quan
CAMELYON17 là dataset mở rộng của CAMELYON16, gồm ảnh WSI hạch bạch huyết canh gác từ 5 trung tâm y tế khác nhau (multi-center), với 1000 WSI (5 slide/bệnh nhân x 200 bệnh nhân). Bài toán chính là phân loại di căn theo 4 mức tại cấp lymph-node (negative/isolated tumor cells/micro-metastases/macro-metastases) và tổng hợp thành pN-stage tại cấp bệnh nhân.
Nguồn dữ liệu: AWS Open Data, s3://camelyon-dataset/CAMELYON17/ (region us-west-2, truy… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON17.cameo_data
CAMEO Dataset for Protein Structure Prediction
This dataset contains protein sequences and structures from CAMEO (Continuous Automated Model EvaluatiOn) for monomer structure prediction tasks.
Dataset Description
CAMEO is a community-wide initiative to continuously evaluate the performance of protein structure prediction methods. This dataset includes 141 protein targets collected from January to April 2025.
Dataset Structure
cameo/
├── sequences.fasta… See the full description on the dataset page: https://huggingface.co/datasets/THU-ATOM/cameo_data.InferenceNetInferenceNet project portal · Home · Data overview · Leaderboard · Agent / Harness
Explore the project and its published research results in the linked Space. This dataset repository remains the source for the task lists and research data.
InferenceNet: Data Card for Econometric AI Agent Testset
InferenceNet is a project aimed at evaluating and building up the AI capability for social science research related to empirical studies. We collect the world’s largest dataset on… See the full description on the dataset page: https://huggingface.co/datasets/CamoAiLab/InferenceNet.Campus_Recruitment_CSV
Dataset Description
This data set consists of Placement data of students in a XYZ campus. Based on the student's performance data we are classifying his Placement Status.
The students report includes the following information:
CGPA - The grade of the student in his university
Internships - The no of internship done by the student before final placement
Projects - The no of projects done by the student
Workshops/Certifications - The no of workshops attended and the certifications… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_CSV.Campus_Recruitment_Text
Dataset Description
This data set consists of Placement data of students in a XYZ campus. Based on the student's performance report we are classifying his Placement Status. The dataset is derived from a csv data.
The Mistral7B model is used with data-to-text methodology to convert each of the rows in the csv data into a textual format for the LLM's, the conversion script is in this notebook.
The Prompt field is the prompt used on Mistral7B LLM and the response field is the… See the full description on the dataset page: https://huggingface.co/datasets/Krooz/Campus_Recruitment_Text.tmax-fair-sc-campaign-curves
TMax fair-sc DPPO training curves
Per-optimizer-step training curves for the fair-sc DPPO arms of the TMax
terminal-RL campaign, exported from the Meta-internal W&B
(meta-fair.wandb.io, entity oscaryinn, project open-instruct-terminal-rl).
Not the SC3 campaign. TMaxxx/tmax-sc3-campaign-curves holds the 13 SC3
arms, exported from oscaryin4422/tmax-sc3-campaign on public wandb.ai, from a
different cluster. The arms here never appear there: fair-sc compute nodes
cannot route to… See the full description on the dataset page: https://huggingface.co/datasets/TMaxxx/tmax-fair-sc-campaign-curves.marketing_campaign_datamarketing_campaigncamerabench_vqa_lmms_evalad_campaign_datasetmarketing-campaign-performance-200k
Marketing Campaign Performance Dataset (200k)
Mirror de Kaggle: Marketing Campaign Performance Dataset (manishabhatt22), 200.000 filas de campañas de marketing (verificado: rango de fechas 2021-01-01 a 2021-12-31, no dos años como dice la card de Kaggle). Subido aquí para poder cargarlo con datasets.load_dataset() sin credenciales de Kaggle.
Contenido
data/marketing_campaign_dataset.csv — 200.000 filas, ~27 MB, 16 columnas.
data/data_dictionary.md — descripción… See the full description on the dataset page: https://huggingface.co/datasets/federicomoreno/marketing-campaign-performance-200k.email-campaigns
Email Marketing Campaign Analytics Dataset (Free Sample)
This is a free sample with 4,025 rows. The full dataset has 56,462 rows across 4 tables.
Email campaign performance data for a simulated B2B SaaS company running
120 campaigns over 18 months. 15,000 subscribers across 5 segments,
40,000 email events (sends, opens, clicks, bounces, unsubscribes).
Features realistic engagement curves: declining open rates over time,
segment-specific behavior, A/B test results, and two… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/email-campaigns.cambioML-QA_v2camie-tagger-vs-wd-tagger-val
What's what
1_json_to_csv.py:
converts cm_tags.json to a csv format I'm more used to and that is easier to use with my existing tooling
2_retrieve_images_by_cc.py:
retrieves the validation images using cheesechaser, downloads them to "original/"
3_common_tags.py:
clean up both models tag sets to only consider the common tags; note down the indexes to use to fetch the correct tag probs from the dumps generated by the inference scripts
4_cm_onnx_inference.py:
run… See the full description on the dataset page: https://huggingface.co/datasets/SmilingWolf/camie-tagger-vs-wd-tagger-val.BAREC-Shared-Task-2025-sent
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-sent.FIGNEWS-2024This dataset contains the data submitted by all teams in the FIGNEWS shared task as part of the the Second Arabic Natural Language Processing Conference (ArabicNLP 2024).
You can find more details about this shared task at the FIGNEWS homepage.
For more details about the data, please check the FIGNEWS github page.
@misc{zaghouani2024fignewssharedtasknews,
title={The FIGNEWS Shared Task on News Media Narratives},
author={Wajdi Zaghouani and Mustafa Jarrar and Nizar Habash and… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/FIGNEWS-2024.ahsan81_superstore-marketing-campaign-dataset
Superstore Marketing Campaign Dataset
Sample customer data for analysis of a targeted Membership Offer
Dataset Info
Source: Kaggle
Original Size: 0.05 MB
Kaggle Downloads: 18,826
Files: 1
Files
superstore_data.csv
Mirrored from Kaggle
fake-email-campaigninsurance-campaign-data
Key Features
Primary Research Focus:
Age vs Income correlation
Occupation, Household size
Marital Status
Dataset Highlights
Dataset shape: (15000, 61)
Number of features: 59
Generation date: 2025-07-20 13:56:12.559153
Random seed: 42
Feature Categories:
Customer Profile: 7 features
Asset Ownership: 6 features
Behavioral: 5 features
Product Portfolio: 5 features
Temporal: 5 features
Data types:
int64 41
float64 20
Name: count… See the full description on the dataset page: https://huggingface.co/datasets/nprak26/insurance-campaign-data.BAREC-Shared-Task-2025-doc
BAREC Shared Task 2025
Dataset Summary
BAREC (the Balanced Arabic Readability Evaluation Corpus) is a large-scale dataset developed for the BAREC Shared Task 2025, focused on fine-grained Arabic readability assessment. The dataset includes over 1M words, annotated across 19 readability levels, with additional mappings to coarser 7, 5, and 3 level schemes.
The dataset is annotated at the sentence level. Document-level readability scores are derived by assigning each… See the full description on the dataset page: https://huggingface.co/datasets/CAMeL-Lab/BAREC-Shared-Task-2025-doc.aml-campaigngraph-data
AML-CampaignGraph Data Card
Dataset summary
AML-CampaignGraph v1.0.0 is a fully synthetic temporal transaction-graph benchmark for campaign-level anti-money-laundering research. It is designed for reproducible experiments on campaign ranking, temporal generalization, controlled out-of-distribution evaluation, and evidence extraction. The benchmark deliberately uses generated data because public banking records contain sensitive financial and personal information… See the full description on the dataset page: https://huggingface.co/datasets/adnanallemon/aml-campaigngraph-data.strange-purpose-60cd2e
strange-purpose-60cd2e
Synthetic products test data: 47 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/camposalexis/strange-purpose-60cd2e.huggingface_9687-780rcq4k-campaign-registry
Campaign Registry (Roy229/huggingface_9687-780rcq4k-campaign-registry)
Shared campaign creative-brief registry used by marketing operations.
Business records live in data/campaigns.csv.
CAMEO
Dataset Card for CAMEO
Dataset to accompany the EMNLP'23 paper titled: "Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection".
Dataset Details
50,000 subset from the GoEmotions Dataset automatically annotated with the following linguistic complexity measures:
idt: Incomplete Dependency Theory
dlt: Dependency Locality Theory
nnd: Nested-Nouns Distance
le: Left-embededness
percentage_polysyllable_words: % of polysyllable words… See the full description on the dataset page: https://huggingface.co/datasets/pranaydeeps/CAMEO.caml_naicsoverall-history-9ccb5b
overall-history-9ccb5b
Synthetic weather test data: 57 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/camposchristopher/overall-history-9ccb5b.
