datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Qwen3-8B-customerservice-LLM-as-a-judge-dataGPT-4.1-customerservice-LLM-as-a-judge-dataQwen3-1.7B-customerservice-LLM-as-a-judge-dataSmolLM3-3B-customerservice-LLM-as-a-judge-dataQwen3-4B-customerservice-LLM-as-a-judge-dataLlama3.2-1b-instruct-customerservice-LLM-as-a-judge-datallm-as-a-judge-eli5gemini-2.5-flash-customerservice-LLM-as-a-judge-datacustomerservice-llm-as-a-judge-task-resultsLlama3.1-8b-instruct-customerservice-LLM-as-a-judge-dataGemma3-4B-instruct-customerservice-LLM-as-a-judge-dataLlama3.2-3B-instruct-customerservice-LLM-as-a-judge-dataPhi-4-mini-customerservice-LLM-as-a-judge-dataVirtuoso-large-customerservice-LLM-as-a-judge-data
Virtuoso-large-customerservice-LLM-as-a-judge-data
Dataset updated with new evaluation columns.
This README refresh triggers Hugging Face metadata re-index.
