datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prometheus-llm-as-a-judge-v1repro-efficient-inference-for-noisy-llm-as-a-judge-evaluation-traces
Agent traces
Agent sessions published from a Trackio Logbook.
dentaku-llm-as-a-judge(English part follows Japanese one.)
team-dentaku/dentaku-llm-as-a-judge
FT-LLM2026チューニングコンペティション (数学タスク) におけるチームdentakuの提出システム構築に利用したデータセットです。
本リポジトリでは正誤判定予測モデルの学習に利用したデータセットを公開しています。
各列名 (キー名) とその内容については以下の通りです
question: 入力となる数学問題
code: questionに対する解法として期待されるコード
decision: codeが与えられた数学の問題文の解答として妥当であると考えられる場合True,妥当でない場合はFalse.
rationale: 最終判断 (decision) に至る理由を説明するテキスト
thinking: 推論過程
正誤判定予測モデルの学習時には入力に入力プロンプトテンプレートを用い,出力に
出力プロンプトテンプレートを利用しています。… See the full description on the dataset page: https://huggingface.co/datasets/team-dentaku/dentaku-llm-as-a-judge.Qwen3-8B-customerservice-LLM-as-a-judge-dataGPT-4.1-customerservice-LLM-as-a-judge-dataQwen3-1.7B-customerservice-LLM-as-a-judge-dataLlama3.1-8b-customerservice-LLM-as-a-judge-dataSmolLM3-3B-customerservice-LLM-as-a-judge-dataQwen3-4B-customerservice-LLM-as-a-judge-dataLlama3.2-1b-instruct-customerservice-LLM-as-a-judge-datallm-as-a-judge-eli5gemini-2.5-flash-customerservice-LLM-as-a-judge-datacustomerservice-llm-as-a-judge-task-resultsLlama3.1-8b-instruct-customerservice-LLM-as-a-judge-dataGemma3-4B-instruct-customerservice-LLM-as-a-judge-dataLlama3.2-3B-instruct-customerservice-LLM-as-a-judge-dataPhi-4-mini-customerservice-LLM-as-a-judge-dataVirtuoso-large-customerservice-LLM-as-a-judge-data
Virtuoso-large-customerservice-LLM-as-a-judge-data
Dataset updated with new evaluation columns.
This README refresh triggers Hugging Face metadata re-index.
