CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lighteval /winograd_wsc Dataset Card for The Winograd Schema Challenge Dataset Summary A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its resolution. The schema takes its name from a well-known example by Terry Winograd: The city councilmen refused the demonstrators a permit because they [feared/advocated] violence. If the… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/winograd_wsc.tabularmultiple-choicen<1K0 likes5.7k downloads1y agoHugging Face02Lansechen /details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval. The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.tabular1K<n<10K0 likes607 downloads1y agoHugging Face03lighteval /hendrycks_ethicstabular100K<n<1M0 likes382 downloads1y agoHugging Face04lighteval /RULER-262144-gemma3-instructtabular1K<n<10K0 likes177 downloads1y agoHugging Face05lighteval /RULER-32768-Qwen-3-Instructtabular1K<n<10K2 likes174 downloads1y agoHugging Face06lighteval /RULER-16384-Qwen-3tabular1K<n<10K0 likes158 downloads1y agoHugging Face07lighteval /RULER-8192-Qwen-3tabular1K<n<10K0 likes144 downloads1y agoHugging Face08lighteval /RULER-16384-Lamma3-Instructtabular1K<n<10K0 likes124 downloads1y agoHugging Face09lighteval /RULER-131072-Lamma3-Instructtabular1K<n<10K0 likes121 downloads1y agoHugging Face10lighteval /RULER-8192-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes121 downloads1y agoHugging Face11lighteval /RULER-4096-Qwen-3tabular1K<n<10K0 likes120 downloads1y agoHugging Face12lighteval /RULER-65536-Falcon-H1-3B-Basetabular1K<n<10K0 likes116 downloads1y agoHugging Face13lighteval /RULER-65536-Qwen-3tabular1K<n<10K0 likes110 downloads1y agoHugging Face14lighteval /RULER-16384-Qwen-3-Instructtabular1K<n<10K0 likes108 downloads1y agoHugging Face15lighteval /RULER-131072-Qwen-3tabular1K<n<10K0 likes101 downloads1y agoHugging Face16lighteval /RULER-65536-Qwen-3-Instructtabular1K<n<10K0 likes95 downloads1y agoHugging Face17lighteval /RULER-65536-Qwen2.5-Instructtabular1K<n<10K0 likes95 downloads1y agoHugging Face18lighteval /RULER-262144-gemma3-basetabular1K<n<10K0 likes93 downloads1y agoHugging Face19juliadollis /energy-mcq-harder-lighteval-rag energy-mcq-harder-lighteval-rag Avaliação do pipeline de RAG da Cemig pelo LightEval, através de um endpoint compatível com OpenAI. O RAG é avaliado como se fosse um modelo: mesmo runner, mesmas tasks e mesmas métricas usadas nos modelos puros. Como ler Uma linha por (experimento, modelo, task, métrica). A coluna model_name é o nome do experimento — decoder-only é a baseline sem recuperação, e rag-<encoder>[-<reranker>] identifica a combinação usada.… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/energy-mcq-harder-lighteval-rag.tabularn<1K0 likes92 downloads1mo agoHugging Face20lighteval /RULER-8192-Qwen-3-Instructtabular1K<n<10K0 likes90 downloads1y agoHugging Face21lighteval /RULER-262144-Qwen-3-Instructtabular1K<n<10K0 likes89 downloads1y agoHugging Face22lighteval /RULER-131072-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes88 downloads1y agoHugging Face23lighteval /RULER-32768-Lamma3-Instructtabular1K<n<10K0 likes86 downloads1y agoHugging Face24lighteval /RULER-8192-Qwen2.5-Instructtabular1K<n<10K0 likes86 downloads1y agoHugging Face25lighteval /RULER-65536-Lamma3-Instructtabular1K<n<10K0 likes85 downloads1y agoHugging Face26lighteval /RULER-32768-Qwen-3tabular1K<n<10K0 likes85 downloads1y agoHugging Face27juliadollis /energy-mcq-harder-lighteval-rag-easy energy-mcq-harder-lighteval-rag-easy Avaliação do pipeline de RAG da Cemig pelo LightEval, através de um endpoint compatível com OpenAI. O RAG é avaliado como se fosse um modelo: mesmo runner, mesmas tasks e mesmas métricas usadas nos modelos puros. Como ler Uma linha por (experimento, modelo, task, métrica). A coluna model_name é o nome do experimento — decoder-only é a baseline sem recuperação, e rag-<encoder>[-<reranker>] identifica a combinação usada.… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/energy-mcq-harder-lighteval-rag-easy.tabularn<1K0 likes84 downloads1mo agoHugging Face28lighteval /RULER-262144-Lamma3-Instructtabular1K<n<10K0 likes77 downloads1y agoHugging Face29lighteval /RULER-262144-SmolLM3-11T-32k-v1-remote-codetabular1K<n<10K0 likes77 downloads1y agoHugging Face30juliadollis /energy-mcq-harder-lighteval-rag-hard energy-mcq-harder-lighteval-rag-hard Avaliação do pipeline de RAG da Cemig pelo LightEval, através de um endpoint compatível com OpenAI. O RAG é avaliado como se fosse um modelo: mesmo runner, mesmas tasks e mesmas métricas usadas nos modelos puros. Como ler Uma linha por (experimento, modelo, task, métrica). A coluna model_name é o nome do experimento — decoder-only é a baseline sem recuperação, e rag-<encoder>[-<reranker>] identifica a combinação usada.… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/energy-mcq-harder-lighteval-rag-hard.tabularn<1K0 likes77 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.