datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sop-bench
SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
📄 Paper: SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents
🏭 Human Expert-Authored SOPs · 🤖 Human-AI Collaborative Framework · 📊 Executable Interfaces · 🔧 Two Agent Architectures · 📈 11 Frontier Models Evaluated
Dataset Summary
SOP-Bench is a comprehensive benchmark for evaluating LLM-based agents on complex, multi-step Standard Operating Procedures (SOPs) that are fundamental to industrial… See the full description on the dataset page: https://huggingface.co/datasets/amazon/sop-bench.xtr-wiki_qa
Xtr-WikiQA
Dataset Summary
Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages.
This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face).
For translations, we used Amazon Translate.
Languages
Arabic (ar)
Spanish (es)
French (fr)
German (de)
Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/xtr-wiki_qa.
