datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchmark-evaluation-resultsrivers-evaluation-results
Rivers Evaluation Results - Comprehensive LLM Benchmarking
All results from the paper's five experimental conditions: baseline LLMs, fine-tuned models, RAG, and Graph-RAG with Licensing Oracle. This repository contains baseline evaluations for Claude Sonnet 4.5, Gemini 2.5 Flash Lite, and Gemma 3-4B, along with fine-tuning results for both factual recall and abstention behavior. It also includes outputs from the embedding-based RAG system and the Graph-RAG with Licensing Oracle… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/rivers-evaluation-results.
