datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
epiq-laer-corporate-benchmark
Epiq LAER CorporateBench
Harbor release of Epiq LAER CorporateBench: Enterprise Knowledge.
Epiq LAER CorporateBench is CorporateBench in its Harbor configuration. It packages the five CorporateBench capabilities as 128 scored Harbor tasks with 1,132 graded cases over the four synthetic companies of the original paper, with five development tasks alongside.
What is this?
LLMs are increasingly able to answer complex questions about enterprise-scale document… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/epiq-laer-corporate-benchmark.corporatebench
CorporateBench
Dataset release for CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases.
CorporateBench evaluates information extraction, retrieval, and question answering over four synthetic corporate corpora ranging from 353 to 232,692 released documents. The corpora are generated from temporally evolving knowledge bases, providing deterministic ground truth across related documents.
Dataset Viewer
https://corporatebench.epiqai.com/… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/corporatebench.corporateDataset
Corporate Data Analysis Training Dataset (Clean)
Dataset Description
This is a cleaned and standardized corporate analysis training dataset with consistent schema.
Schema
All entries follow the instruction-input-output format:
{
"instruction": "Task description",
"input": "Business data or context",
"output": "Analysis and insights"
}
Features
✅ Consistent Schema - All entries use the same format
✅ Clean Data - Validated and error-free
✅… See the full description on the dataset page: https://huggingface.co/datasets/MikePfunk28/corporateDataset.corporate-site-harness-training-data
Dataset Card for Corporate Site Harness Training Data
Revision: v0.3-lora-standard
Factory git commit: ad72ba654590dd83d7070b23990b164ffae688a4
Dataset Summary
English chat-style supervised fine-tuning (SFT), preference (DPO), and held-out
evaluation data for teaching a local LLM the corporate/site harness used by
corporate-site-harness:
policy — phases, roles, workspace isolation, premium-model routing, factory vs product
cli — corp-harness argv, tool-grounded… See the full description on the dataset page: https://huggingface.co/datasets/SafetyMP/corporate-site-harness-training-data.swahili_corporate_rag_i
Swahili Corporate RAG I
A 10K-entry Supervised Fine-Tuning (SFT) / RAG dataset in Swahili, generated using Gemini 3.5.
Designed specifically for training enterprise assistants to understand corporate context, policies, and customer service instructions in Swahili.
