datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
polish-tax-interpretations
Polish Tax Interpretations (Eureka)
A corpus of 538,866 Polish tax-law interpretations and binding rulings published by the
Polish Ministry of Finance / Krajowa Informacja Skarbowa (KIS) on the public Eureka portal
(eureka.mf.gov.pl), together with their official portal metadata.
Each row is one ruling: the full text body plus the scraped official metadata (signature,
headnote, issuing authority, dates, categories, keywords, cited legal provisions, classifications).
Curated by… See the full description on the dataset page: https://huggingface.co/datasets/czlonkowski/polish-tax-interpretations.interpretation-kr
wellsa-ai interpretation-kr
법령해석례 (statutory interpretations).
MiniLex 7-domain Korean lawdata infrastructure.
Snapshot
Documents: 17,354
Snapshot date: 2026-06-05
Source: 법제처 DRF OpenAPI
Pipeline: daily cron 06:15~07:30 KST (scrape → fetch_body → convert → commit)
Schema
Each row in train.jsonl:
field
type
description
id
string
stable document id
name
string
document name (Korean)
category
string
subdirectory (year / type / dept)… See the full description on the dataset page: https://huggingface.co/datasets/wellsa-ai/interpretation-kr.structured_poem_interpretation_corpus_staleMasking policy: For rows with source == "poetry_foundation", the poem and
interpretation fields are set to null to respect content licensing. Public-domain
entries (source == "public_domain_poetry") include full text. All categorical annotations
(emotions, primary_emotion, sentiment, themes, themes_50) and metadata remain available.
Tool_Output_Interpretation_Normalization
🇰🇿 Kazakh Tool Output Interpretation and Financial Action Dataset
Dataset Summary
Kazakh Tool Output Interpretation and Financial Action Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in tool-augmented agentic workflows that require interpreting structured tool outputs and generating grounded final responses.
The dataset focuses on scenarios where an assistant must understand a Kazakh user request, call the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Output_Interpretation_Normalization.UN_NU_interpretation_LLMs
Quantifier Scope Interpretation Dataset
Datasets for an ongoing project about Scope preferences and ambiguity in LLM interpretation.
Dataset Structure
Splits
The dataset consists of synthetically generated stimuli pairing target sentences with interpretation-biased contexts (SSR vs. ISR).
Features
language (string)Language of the stimulus (English or Chinese).
structure (string)Surface syntactic configuration of the sentence:UN (universal >… See the full description on the dataset page: https://huggingface.co/datasets/CALM-Lab-Purdue/UN_NU_interpretation_LLMs.
