datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CRMArenaPro
Dataset Card for CRMArena-Pro
Dataset Description
Paper Information
Citation
Dataset Description
CRMArena-Pro is a benchmark for evaluating LLM agents' ability to perform real-world work tasks in realistic environment. It expands on CRMArena with nineteen expert-validated tasks across sales, service, and "configure, price, and quote" (CPQ) processes, for both Business-to-Business (B2B) and Business-to-Customer (B2C) scenarios. CRMArena-Pro distinctively incorporates… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/CRMArenaPro.CRMArena
Dataset Card for CRMArena
Dataset Description
Paper Information
Citation
Dataset Description
CRMArena is a benchmark for evaluating LLM agents' ability to perform real-world work tasks in realistic environment. This benchmark is introduced in the paper "CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments". We include 16 commonly-used industrial objects (e.g., account, order, knowledge article, case) with… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/CRMArena.synthetic-crm-seed-data-for-property-management
Synthetic CRM Seed Data for Concierge & Property Management: GDPR-compliant, realistic market distributions for software testing & demos (SQL/CSV) (DEMO)
Thank you for downloading this synthetic dataset!
This repository contains a functional, lightweight preview of a 100% synthetic dataset designed for testing real estate CRMs and property management software.
>> Full version available on Gumroad <<
This demo provides a baseline schema. For large-scale stress… See the full description on the dataset page: https://huggingface.co/datasets/Pandoxyd/synthetic-crm-seed-data-for-property-management.sglang-test-backupsScott-Schmitz-Real-Estate-CRM-Secretssglang-migration-20260911
Private migration backup
Code, worktrees, original corpora, experiment outputs and frozen evidence for sglang_test.
Read RESTORE.md and verification.json before restoring. migration-inventory.json maps the included files; SHA256SUMS uses portable relative paths. Research-secure remote data is deferred at the owner request. Runtime KV stores and credentials are excluded.
GitHub source: https://github.com/crmsndu/sglang_test
crmsc-envswmo-crmarena-traces
crmarena — real agent-environment traces
Professional CRM analytics over a realistic Salesforce org snapshot: case routing, handle-time analytics, and entity disambiguation via SQL.
Every trace is a REAL run: an LLM agent stepping against the actual benchmark environment, with
each transition (tool call → true environment observation) recorded as OpenTelemetry GenAI spans
(traces.otel.jsonl, one span per line). Captured by
world-model-harness's
environment-capture package, which… See the full description on the dataset page: https://huggingface.co/datasets/experiential-labs/wmo-crmarena-traces.arc-crm-6
Blobfish Arc CRM 6 · 0.1.2
Six independently authored synthetic CRM workflows with 31 domain/evidence tools
on three mock servers, two engine controls, and 33 synthetic PDF/XLSX/EML assets.
Each episode shares one stateful world across CLI, real HTML forms, REST and MCP.
Scope and results
All six frozen reference solutions pass local Docker execution with strict score
100 and exact final-state/trace parity. All 167 authoring negative controls reject.
These are… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/arc-crm-6.arc-crm-benchmark
Arc CRM Benchmark Dataset
Dataset Description
The Arc CRM Benchmark is a production-realistic synthetic CRM environment dataset for evaluating LLM agents on state-modifying workflows. This dataset provides a comprehensive testbed for measuring agent performance, reliability, and adaptation through continual learning frameworks.
The dataset contains 1,200 multi-turn conversations covering diverse CRM workflows with varying complexity. Each conversation simulates realistic… See the full description on the dataset page: https://huggingface.co/datasets/Arc-Intelligence/arc-crm-benchmark.cidoc-crm-corpus
CIDOC CRM SIG corpus — derived artifacts
The built corpus for cidoc-crm-mcp:
26 years of CIDOC CRM Special Interest Group mailing list, its issue
register and meeting minutes, cleaned, threaded, and indexed for BM25 and
vector search.
These are derived artifacts, not the source. They exist because the code
repository cannot carry them: ~876MB, git-ignored, rebuilt from a 143MB mbox
that is distributed separately.
Fetching
uv run python build.py fetch… See the full description on the dataset page: https://huggingface.co/datasets/stefdoerr/cidoc-crm-corpus.putnam-api-38177-decode-heavyhuggingface_4691_crm_marketing_campaigns
Marketing Campaigns
Snapshot of GLOBAL data.
turkey-business-digital-presence
Turkey Business Digital Presence 2026
Website, phone, email and social media presence for 1,786,700 businesses in Turkey,
broken down by all 81 provinces and 40 sectors.
Full report: English ·
Türkçe ·
DOI: 10.5281/zenodo.22217770 ·
Source repo: CRM-Solid/turkey-business-digital-data
Headline findings
Measure
Share
Count
Website on record
38.3%
684,496
Website that actually answers
23.2%
413,777
Phone number
65.3%
1,166,881
Email address
29.9%
534… See the full description on the dataset page: https://huggingface.co/datasets/crmsolid/turkey-business-digital-presence.CRMpert
CRMpert Dataset
This dataset contains flow fields around wings perturbed from the Common Research Model. It contains 288 shapes and 2145 flow fields, acting as a task-specific
dataset to fine-tune a pre-trained model for a more precise local surrogate model.
Perturbation
The wing geometry is controlled at seven spanwise sections. At each control section, the sectional airfoil is parameterized using 20 CST coefficients, each perturbed within
$\pm 40%$ of the… See the full description on the dataset page: https://huggingface.co/datasets/thuerey-group/CRMpert.COIG-P-CRMThis repository contains the COIG-P-CRM dataset used for the paper COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values.
CRM
📊 B2B CRM Sales Opportunities Dataset
🎯 Dataset Overview
This dataset contains simulated B2B sales pipeline data for a computer hardware company across 5 relational tables:
sales_pipeline: Sales opportunity transaction logs (deal stage, close value, dates).
accounts: B2B customer profiles, employee counts, revenue, and sectors.
products: Product series, hardware items, and suggested retail prices.
sales_teams: Sales agents, managers, and regional office… See the full description on the dataset page: https://huggingface.co/datasets/Nouramoh/CRM.huggingface_4691_crm_customers_us
Customer Profiles US
Snapshot of US data.
lex_crminal_lawzarn-crm-note-to-followup-messages
Zarn CRM Note to Followup Messages
Dataset Description
High-context CRM notes transformed into personalized follow-up emails, DMs, and next-step nudges with factual preservation checks.
Team Attribution
This dataset was created and reviewed by the Zarnite team through internal benchmark design, generation, and quality-control workflows. It should be presented as a Zarnite-authored benchmark starter pack, not as a purely human-collected field corpus.… See the full description on the dataset page: https://huggingface.co/datasets/zarnite/zarn-crm-note-to-followup-messages.crm_dataset_originalhuggingface_4691_crm_ordershuggingface_4691_crm_products
Product Catalog
Snapshot of GLOBAL data.
crm_function_calling_datasethuggingface_4691_crm_customers_eu
Customer Profiles EU
Snapshot of EU data.
CRM_WORKSHOPcustom_crm_datasetcrm-cleaning-dataset-v1synthetic_customer_support_crm
Synthetic Customer Support CRM Dataset
This dataset is a fully synthetic, multi-table relational dataset designed for
experiments in customer-support analytics, SQL training, RAG pipelines,
agent-based reasoning, and machine learning workflows.
All data was generated using ChatGPT and does not contain any real personal
information. Names, emails, phone numbers, and ticket details are entirely
fictional.
📦 Dataset Contents
The dataset includes five CSV files:… See the full description on the dataset page: https://huggingface.co/datasets/air5978/synthetic_customer_support_crm.crm_function_calling_dataset
