datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
greek-bar-bench
Dataset Card for GreekBarBench 🇬🇷🏛️⚖️
GreekBarBench is a benchmark designed to evaluate LLMs on challenging legal reasoning questions across five different legal areas from the Greek Bar exams, requiring citations to statutory articles and case facts.
This repository hosts two related benchmarks:
Benchmark
Subsets
Task
GreekBarBench (GBB)
greekbarbench, gbb-jme
Free-text legal reasoning with citations, and LLM-judge meta-evaluation
GreekBarRetrieval (GBR)… See the full description on the dataset page: https://huggingface.co/datasets/AUEB-NLP/greek-bar-bench.gretel-math-gsm8k-v0
gretelai/gsm8k-synthetic-diverse-405b
This dataset is a synthetically generated version inspired by the GSM8K https://huggingface.co/datasets/openai/gsm8k dataset, created entirely using Gretel Navigator with meta-llama/Meta-Llama-3.1-405B as the agent LLM. It contains ~1500 Grade School-level math word problems with step-by-step solutions, focusing on age group, difficulty, and domain diversity.
Key Features:
Synthetically Generated: Math problems created using Gretel… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-math-gsm8k-v0.synthetic_multilingual_llm_prompts
Image generated by DALL-E. See prompt for more details
📝🌐 Synthetic Multilingual LLM Prompts
Welcome to the "Synthetic Multilingual LLM Prompts" dataset! This comprehensive collection features 1,250 synthetic LLM prompts generated using Gretel Navigator, available in seven different languages. To ensure accuracy and diversity in prompts, and translation quality and consistency across the different languages, we employed Gretel Navigator both as a generation tool and as an… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_multilingual_llm_prompts.OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
Dataset Description
OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation.
Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.emission-factor-benchmark
Emission-Factor Accuracy Benchmark
3,299 rows. Five frontier models answering identical factual questions, with
ground truth traced to a named document and an exact cell — plus the same
questions re-run with a lookup tool, and a second study on which data vendors
those models recommend unprompted.
Collected 10 September 2026. Models: claude-opus-5, gpt-5.5,
gemini-3.1-pro-preview, gemini-3.6-flash, grok-4.6. All answers were
produced through each provider's API with no tools and… See the full description on the dataset page: https://huggingface.co/datasets/greencalculus/emission-factor-benchmark.MMLU-Pro_greekgreek_pcr
Dataset Card for Greek physical commonsense reasoning
The GPCR (Greek physical commonsense reasoning) resource is a manually-annotated physical commonsense reasoning evaluation dataset for Greek.
GPRC includes 208 PIQA-style examples consisting of a prompt with two candidate completions ("solutions").
The "solution" pairs are as similar as possible, only differing in one or two words in most of the cases.
Approximately 40% of the examples are culturally specific, i.e. they… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/greek_pcr.
