corporate
legalbench_corporate_lobbying
LegalBenchCorporateLobbying
An MTEB dataset
Massive Text Embedding Benchmark
The dataset includes bill titles and bill summaries related to corporate lobbying.
Task category
t2t
Domains
Legal, Written
Reference
https://huggingface.co/datasets/nguha/legalbench/viewer/corporate_lobbying
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/legalbench_corporate_lobbying.epiq-laer-corporate-benchmark
Epiq LAER CorporateBench
Harbor release of Epiq LAER CorporateBench: Enterprise Knowledge.
Epiq LAER CorporateBench is CorporateBench in its Harbor configuration. It packages the five CorporateBench capabilities as 128 scored Harbor tasks with 1,132 graded cases over the four synthetic companies of the original paper, with five development tasks alongside.
What is this?
LLMs are increasingly able to answer complex questions about enterprise-scale document… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/epiq-laer-corporate-benchmark.entity-references
Entity References Database
A comprehensive entity database for organizations, people, roles, and locations with embedding-based semantic search. Built from authoritative sources (GLEIF, SEC, Companies House, Wikidata) for entity linking and named entity disambiguation.
Dataset Summary
This dataset provides fast lookup and qualification of named entities using vector similarity search. It stores records from authoritative global sources with embeddings generated by… See the full description on the dataset page: https://huggingface.co/datasets/Corp-o-Rate-Community/entity-references.corporatebench
CorporateBench
Dataset release for CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases.
CorporateBench evaluates information extraction, retrieval, and question answering over four synthetic corporate corpora ranging from 353 to 232,692 released documents. The corpora are generated from temporally evolving knowledge bases, providing deterministic ground truth across related documents.
Dataset Viewer
https://corporatebench.epiqai.com/… See the full description on the dataset page: https://huggingface.co/datasets/epiq-ai-labs/corporatebench.corporate-actions
US Corporate Actions — dividends and splits
391 639 dividends from 3 327 filers · 5 619 splits from 3 814 filers ·
2005 to 2026
Built to close a specific hole. A filing states shares and earnings per share
as of the day it was made; every price series is adjusted for splits since.
Multiply one by the other and the answer is wrong by the split factor — on
Deckers that turned a 6.9% earnings yield into 41.7%, a P/E of 1.8.
The pipeline lives in recipe/ at the same revision as the… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/corporate-actions.corporate-emission-reports
Dataset Card for Dataset Name
A dataset of 100 corporate sustainability reports with manually extracted scope 1, 2 and 3 greenhouse gas emission values.
Dataset Details
Dataset Description
Data about corporate greenhouse gas emissions is usually published only as part of sustainability report PDF's, which is not a machine-readable format. Interested actors have to manually extract emission data from these reports, which is a tedious and time-consuming process.… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/corporate-emission-reports.
