synthetic-evaluation
synthetic-llm-evaluation-traces
EvaluLLM-Inspired Synthetic Evaluation Traces (DBbun)
View source code on GitHub
Watch on Youtube: Evaluating AI with AI
Dataset Summary
This dataset contains fully synthetic evaluation traces for NLG / LLM-style output comparison, inspired by the evaluation workflow described in EvaluLLM: LLM Assisted Evaluation of Generative Outputs (IUI Companion 2024).
The dataset is produced by a configurable, offline simulator and includes:
synthetic tasks (prompts)
synthetic… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/synthetic-llm-evaluation-traces.natural_language_prompt_synthetic_dataset_evaluation_instruct_datasetllama_all_synthetic_dataset_evaluation_instruct_datasetnatural_language_prompt_w_correct_ans_synthetic_dataset_evaluation_instruct_datasetnatural_language_prompt_w_correct_ans_synthetic_dataset_evaluation_json_instruct_dataset
