datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DataClawEval
DataClawEval
An executable benchmark for end-to-end data-engineering agents in industrial environments.
DataClawEval measures an autonomous agent's ability to inspect data, implement and debug pipelines,
and materialize correct artifacts in realistic data-engineering workflows. It contains 100
production-grounded tasks across five execution engines: PySpark, MySQL, HiveSQL, PrestoSQL/Trino,
and FlinkSQL. Each task runs in an isolated Docker sandbox and is evaluated by a… See the full description on the dataset page: https://huggingface.co/datasets/dicemy/DataClawEval.DICE-BENCH
🎲 DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
🔗 Links for Reference
Repository: https://github.com/snuhcc/DICE-Bench
Paper: https://arxiv.org/abs/2506.22853
Project page: https://snuhcc.github.io/DICE-Bench/
Point of Contact: kyochul@snu.ac.kr
📖 Paper Description
DICE-BENCH is a benchmark that tests how well large language models can call external functions in realistic… See the full description on the dataset page: https://huggingface.co/datasets/OfficerChul/DICE-BENCH.
