CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /legalbench_corporate_lobbying LegalBenchCorporateLobbying An MTEB dataset Massive Text Embedding Benchmark The dataset includes bill titles and bill summaries related to corporate lobbying. Task category t2t Domains Legal, Written Reference https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.texttext-retrievaln<1K0 likes895 downloads7mo agoHugging Face02lbrenap1 /mining-legal-arguments-us-corporate-case-law Mining Legal Arguments in U.S. Corporate Case Law This dataset contains span-level functional labels and directed support relations for 42 U.S. federal tax opinions concerning corporate reorganizations under I.R.C. Section 368. The opinions range in citation year from 1935 to 1987. Two law students annotated the cases, and a law professor adjudicated the final case-level representations. Ten cases also include the two independent annotations used for inter-annotator agreement… See the full description on the dataset page: https://huggingface.co/datasets/lbrenap1/mining-legal-arguments-us-corporate-case-law.tabulartext-classification10K<n<100K0 likes90 downloads24d agoHugging Face03MikePfunk28 /corporateDataset Corporate Data Analysis Training Dataset (Clean) Dataset Description This is a cleaned and standardized corporate analysis training dataset with consistent schema. Schema All entries follow the instruction-input-output format: { "instruction": "Task description", "input": "Business data or context", "output": "Analysis and insights" } Features ✅ Consistent Schema - All entries use the same format ✅ Clean Data - Validated and error-free ✅… See the full description on the dataset page: https://huggingface.co/datasets/MikePfunk28/corporateDataset.textquestion-answering10K<n<100K1 likes53 downloads1y agoHugging Face04garvitupdy /Corporate_AI_Datasettexttext-generation1K<n<10K0 likes38 downloads26d agoHugging Face05LukeIrwin /corporate-governance-reasoning Dataset Summary The corporate-governance-reasoning dataset was designed to test a model's ability to reason about executive/board/shareholder proposals to alter companies' corporate governance structures. While there are multiple legal datasets, none are focused specifically on reasoning tasks. On the other hand, reasoning (ratiocination and ability to make connections to precedent) is a core part of the practice of the law in the real world. We focus on a specific aspect of the… See the full description on the dataset page: https://huggingface.co/datasets/LukeIrwin/corporate-governance-reasoning.textn<1K4 likes30 downloads1y agoHugging Face06SafetyMP /corporate-site-harness-training-data Dataset Card for Corporate Site Harness Training Data Revision: v0.3-lora-standard Factory git commit: ad72ba654590dd83d7070b23990b164ffae688a4 Dataset Summary English chat-style supervised fine-tuning (SFT), preference (DPO), and held-out evaluation data for teaching a local LLM the corporate/site harness used by corporate-site-harness: policy — phases, roles, workspace isolation, premium-model routing, factory vs product cli — corp-harness argv, tool-grounded… See the full description on the dataset page: https://huggingface.co/datasets/SafetyMP/corporate-site-harness-training-data.texttext-generation1K<n<10K0 likes25 downloads1mo agoHugging Face07anhnon /vietnamese-corporate-legal-articles-fsm Lexora Knowledge - Vietnamese Legal Documents Dataset Summary A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc. Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.texttext-generation100K<n<1M1 likes23 downloads2mo agoHugging Face08axondendriteplus /IC38-Corporate-Agent-Questionstext1K<n<10K0 likes11 downloads1y agoHugging Face09AlexLeoTz /swahili_corporate_rag_i Swahili Corporate RAG I A 10K-entry Supervised Fine-Tuning (SFT) / RAG dataset in Swahili, generated using Gemini 3.5. Designed specifically for training enterprise assistants to understand corporate context, policies, and customer service instructions in Swahili. texttext-generation1K<n<10K0 likes11 downloads3mo agoHugging Face10fingriffin /natural-questions-corporate-jargontextn<1K1 likes3 downloads6mo agoHugging Face11preetu098 /corporatespacetextn<1K0 likes2 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.