CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JesseLiu /patient-evaluations Patient Evaluations Dataset This dataset contains clinician evaluations of AI-generated patient summaries from MIMIC-III data. Dataset Description The dataset includes expert clinician assessments of AI-generated patient summaries, with detailed ratings across multiple dimensions including clinical accuracy, completeness, relevance, and identification of hallucinations or critical omissions. Dataset Structure The dataset contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/JesseLiu/patient-evaluations.texttext-generationn<1K0 likes631 downloads7mo agoHugging Face02furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes101 downloads18d agoHugging Face03Chemin-AI /advent_of_code_evaluations Advent of Code Evaluation This evaluation is conducted on the advent of code dataset on several models including Qwen2.5-Coder-32B-Instruct, DeepSeek-V3-fp8, Llama-3.3-70B-Instruct, GPT-4o-mini, DeepSeek-R1.The aim is to to see how well these models can handle real-world puzzle prompts, generate correct Python code, and ultimately shed light on which LLM truly excels at reasoning and problem-solving.We used pass@1 to measure the functional correctness. Results… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/advent_of_code_evaluations.texttext-generationn<1K2 likes51 downloads2y agoHugging Face043RAIN /brand-bias-evaluations Brand Bias in LLM Recommendations Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains. Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization Quick Start from datasets import load_dataset # Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.tabulartext-generation10K<n<100K0 likes24 downloads6mo agoHugging Face05toolazyhhh123 /hellora-olmoe-gsm8k-evaluations-seed42 OLMoE HELLoRA replication: held-out GSM8K generations Reference archive containing every held-out GSM8K generation used to compare the pinned pretrained base, selective HELLoRA, and full LoRA in the single-GPU replication. Source: https://github.com/toolanzyhhh1234/HELLoRA-replication HELLoRA checkpoint: https://huggingface.co/toolazyhhh123/hellora-olmoe-1b-7b-gsm8k-seed42 Full LoRA checkpoint: https://huggingface.co/toolazyhhh123/lora-olmoe-1b-7b-gsm8k-seed42… See the full description on the dataset page: https://huggingface.co/datasets/toolazyhhh123/hellora-olmoe-gsm8k-evaluations-seed42.text-generation0 likes18 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.