datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
winograd_wsc
Dataset Card for The Winograd Schema Challenge
Dataset Summary
A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is
resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its
resolution. The schema takes its name from a well-known example by Terry Winograd:
The city councilmen refused the demonstrators a permit because they [feared/advocated] violence.
If the… See the full description on the dataset page: https://huggingface.co/datasets/lighteval/winograd_wsc.details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset Card for Evaluation run of Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval
Dataset automatically created during the evaluation run of model Lansechen/Qwen2.5-7B-Open-R1-GRPO-math-lighteval.
The dataset is composed of 5 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 23 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/Lansechen/details_Lansechen__Qwen2.5-7B-Open-R1-GRPO-math-lighteval.hendrycks_ethicsRULER-262144-gemma3-instructRULER-32768-Qwen-3-InstructRULER-16384-Qwen-3RULER-8192-Qwen-3RULER-16384-Lamma3-InstructRULER-131072-Lamma3-InstructRULER-8192-SmolLM3-11T-32k-v1-remote-codeRULER-4096-Qwen-3RULER-65536-Falcon-H1-3B-BaseRULER-65536-Qwen-3RULER-16384-Qwen-3-InstructRULER-131072-Qwen-3RULER-65536-Qwen-3-InstructRULER-65536-Qwen2.5-InstructRULER-262144-gemma3-baseenergy-mcq-harder-lighteval-rag
energy-mcq-harder-lighteval-rag
Avaliação do pipeline de RAG da Cemig pelo LightEval, através de um
endpoint compatível com OpenAI. O RAG é avaliado como se fosse um modelo:
mesmo runner, mesmas tasks e mesmas métricas usadas nos modelos puros.
Como ler
Uma linha por (experimento, modelo, task, métrica). A coluna model_name
é o nome do experimento — decoder-only é a baseline sem recuperação, e
rag-<encoder>[-<reranker>] identifica a combinação usada.… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/energy-mcq-harder-lighteval-rag.RULER-8192-Qwen-3-InstructRULER-262144-Qwen-3-InstructRULER-131072-SmolLM3-11T-32k-v1-remote-codeRULER-32768-Lamma3-InstructRULER-8192-Qwen2.5-InstructRULER-65536-Lamma3-InstructRULER-32768-Qwen-3energy-mcq-harder-lighteval-rag-easy
energy-mcq-harder-lighteval-rag-easy
Avaliação do pipeline de RAG da Cemig pelo LightEval, através de um
endpoint compatível com OpenAI. O RAG é avaliado como se fosse um modelo:
mesmo runner, mesmas tasks e mesmas métricas usadas nos modelos puros.
Como ler
Uma linha por (experimento, modelo, task, métrica). A coluna model_name
é o nome do experimento — decoder-only é a baseline sem recuperação, e
rag-<encoder>[-<reranker>] identifica a combinação usada.… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/energy-mcq-harder-lighteval-rag-easy.RULER-262144-Lamma3-InstructRULER-262144-SmolLM3-11T-32k-v1-remote-codeenergy-mcq-harder-lighteval-rag-hard
energy-mcq-harder-lighteval-rag-hard
Avaliação do pipeline de RAG da Cemig pelo LightEval, através de um
endpoint compatível com OpenAI. O RAG é avaliado como se fosse um modelo:
mesmo runner, mesmas tasks e mesmas métricas usadas nos modelos puros.
Como ler
Uma linha por (experimento, modelo, task, métrica). A coluna model_name
é o nome do experimento — decoder-only é a baseline sem recuperação, e
rag-<encoder>[-<reranker>] identifica a combinação usada.… See the full description on the dataset page: https://huggingface.co/datasets/juliadollis/energy-mcq-harder-lighteval-rag-hard.
