Eyerf/agentblackbox-rag-repair-outcomes
AgentBlackBox RAG Repair Outcome Dataset This dataset contains replay-labeled repair outcome data for AgentBlackBox, a counterfactual debugging framework for language agents. The data is built around failed RAG/document-recall agent traces, candidate repairs, counterfactual replay labels, and repair-ranking evaluation outputs. Contents datasets/ world_model_ranker_dataset_v2_train10k/ pointwise/ listwise/ stats.json… See the full description on the dataset page: https://huggingface.co/datasets/Eyerf/agentblackbox-rag-repair-outcomes.
039
