datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self-reflective-apis
Self-Reflective APIs Benchmark
Dataset accompanying the paper "Self-Reflective APIs: Enhancing AI Agent Efficiency Through Structured Semantic Feedback" (Canedo, Grama — Siemens DI SW).
Overview
This dataset contains the benchmark tasks, experiment results, and tidy analysis table used to produce every table and figure in the paper. It covers two experimental domains (recipe conversion and billing/refund policy) and three LLM conditions across adversarial… See the full description on the dataset page: https://huggingface.co/datasets/arquicanedo/self-reflective-apis.SO-Python_QA-API_Usage-tanh_score
Stack Overflow Python Q&A Dataset
Description
Filtered Python Q&A with API_Usage subcategory without:
Images
Links
Blocks of code
Scores in Q1-Q3 scaled with MaxAbsScaler. Tanh function applyed to joint Scores.
