CoolFace
Apppublic

mariaheloysa-ai/rec-llm-evaluation-audit

sourceHugging Facemitupdated 1mo agoView on Hugging Face
1likes
App README

REC, LLM Evaluation Audit

Interactive static portfolio companion to REC, a tested Python pipeline for comparing human and AI evaluations of language-model responses.

The browser-only demo uses a six-row synthetic sample or a compatible uploaded CSV. It calculates evaluator-agreement metrics, applies the public Version 1 decision policy, creates an explainable inspection queue, and exports evaluated results locally.

The sample data is synthetic and illustrative. It is not a benchmark and does not establish real-world model performance. The complete repository contains the full Python validation and reporting pipeline, methodology documentation, CI, and 80 automated tests.