datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aplikacje-prawnicze-mcq
Polish Legal Apprenticeship Entrance Exams — adwokacka/radcowska, notarialna, komornicza (2007–2025)
Native-Polish, single-choice (A/B/C) legal MCQ benchmark built from the official entrance
examinations for the Polish legal apprenticeships, published by the Ministry of Justice:
adwokacka + radcowska (advocate + legal counsel — a single shared test from 2009 on;
two separate exams in 2007),
notarialna (notary),
komornicza (court-enforcement officer / bailiff).
Each item… See the full description on the dataset page: https://huggingface.co/datasets/bartoszkobylinski1/aplikacje-prawnicze-mcq.D4J-Repair
Dataset Summary
D4J-Repair is a curated subset of Defects4J, containing 371 single-function Java bugs from real-world projects. Each example includes a buggy implementation, its corresponding fixed version, and unit tests for verification.
Supported Tasks
Program Repair: Fixing bugs in Java functions
Code Generation: Generating correct implementations from buggy code
Dataset Structure
Each row contains:
task_id: Unique identifier for the task (in format:… See the full description on the dataset page: https://huggingface.co/datasets/barty/D4J-Repair.lit2vec-tldr-bart-dataset
Lit2Vec TL;DR Chemistry Dataset
Summary
The Lit2Vec TL;DR Chemistry Dataset is a curated collection of 19,992 chemistry research abstracts paired with short, TL;DR-style abstractive summaries.It was created to support research in scientific text summarization, semantic indexing, and domain-specific knowledge graph construction.
Unlike generic summarization datasets, this corpus is:
Legally reusable → all abstracts are sourced from CC-BY licensed publications… See the full description on the dataset page: https://huggingface.co/datasets/Bocklitz-Lab/lit2vec-tldr-bart-dataset.
