datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AssetOpsBench
AssetOpsBench
AssetOpsBench is a specialized benchmark designed for evaluating Large Language Models (LLMs) and Multi-Agent systems in industrial operations. It focuses on the intersection of sensor data interpretation, maintenance logic, and Prognostics and Health Management (PHM).
The benchmark enables researchers to test how effectively AI agents can manage complex industrial assets, such as compressors and hydraulic pumps, by applying rule-based logic and diagnostic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/AssetOpsBench.FailureSensorIQ
FailureSensorIQ Dataset
FailureSensorIQ is a Multi-Choice QA (MCQA) dataset that explores the relationships between sensors and failure modes for 10 industrial assets.
|Github | 🏆Leaderboard | 📖Paper |
Dataset Summary
FailureSensorIQ is a Multi-Choice QA (MCQA) dataset that explores the relationships between sensors and failure modes for 10 industrial assets. By only leveraging the information found in ISO documents, we developed a data generation pipeline that… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/FailureSensorIQ.SQL-API-Bench
Dataset Card for Dataset Name
This dataset contains QA that requires DB and API access at the same time. It is composed of two new benchmarks consisting of questions whose answers require a
combination of database and API calls, both of
which are augmentations of the popular Spider
dataset and benchmark.
Benchmark I replaces a fraction of the real Spider database tables with
equivalents that are executed via APIs. This allows us to directly test the mechanism by which
database and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/SQL-API-Bench.POBS
Preference, Opinion, and Belief Survey (POBS)
POBS is a dataset of survey questions designed to uncover preferences, opinions, and beliefs on societal issues.
Each row represents a question with its topic, options, and polarity.
Columns:
topic: Question topic
category: Category
question_id: Unique question ID
question: Survey question text
options: List of possible answers
options_polarity: Numeric polarity for each option (where applicable)
POBS: Preference, Opinion… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/POBS.
