mariaheloysa-ai/rec-llm-evaluation-audit
1
REC, LLM Evaluation Audit
Interactive static portfolio companion to REC, a tested Python pipeline for comparing human and AI evaluations of language-model responses.
The browser-only demo uses a six-row synthetic sample or a compatible uploaded CSV. It calculates evaluator-agreement metrics, applies the public Version 1 decision policy, creates an explainable inspection queue, and exports evaluated results locally.
The sample data is synthetic and illustrative. It is not a benchmark and does not establish real-world model performance. The complete repository contains the full Python validation and reporting pipeline, methodology documentation, CI, and 80 automated tests.
