datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
changelogs-agentic-rag-3docs-generationschangelogs-agentic-rag-10docs-generationspersonalization-agentic-rag-10docs-generationsclog-eval-generations
clog-eval-generations
Unified eval generations from the continual-internalization / code-changelog benchmark suite. Every row is one model trial on one (mode, library, question) cell.
390,800 rows • 83 eval models • 4 modes (DA, CR, RR, IR)
8 trials per cell • sampling: T=0.7, top_p=0.95, top_k=20
Reconstructed prompts (prompt_system / prompt_user) are included so you can see the chat template used. Code snippets and library corpora are stubbed (e.g. <<CODE SNIPPET MASKED>>) to… See the full description on the dataset page: https://huggingface.co/datasets/continual-internalization/clog-eval-generations.personalization-agentic-rag-5docs-generationschangelogs-agentic-rag-5docs-generationspersonalization-agentic-rag-3docs-generations
