datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
goldensets
LEGEX Goldensets: Expert-Coded Review-Table Annotations
This repository contains the expert-coded gold annotations for the LEGEX
benchmark of civil-judgment review-table extraction. 1,548 judgments across
19 jurisdictions have been annotated by hand against a shared 14-field schema
covering monetary outcomes, cost allocation, party structure, and industry
classification. Including independent secondary re-annotations, the release
holds 1,974 annotation rows.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/legexbenchmark/goldensets.NOBLE_GoldenSet-KR_Verified-v3.2.1
NOBLE v3.2.1 GoldenSet (KR, Verified)
한국어 대화 샘플 63개(JSONL)로 구성된 골든셋(검수 통과)입니다.
목표는 짧고 읽히는 톤, 맥락 고정, 안전한 방향 제시, 과부하 줄이기입니다.
Files
data/NOBLE_v3.2.1_GoldenSet_KR_Verified.jsonl — 검수 통과(qa_ok=YES)만 모은 샘플
SCENARIO_INDEX.md — 질문(상황) 목록 + 태그
Record format (one line = one JSON)
각 줄은 1개의 레코드이며, 주요 필드는 아래와 같습니다.
scenario : 질문/상황(한 문장 요약)
tags : 분류 태그(복수)
signals : (선택) 위험/왜곡 신호 요약
model_response : 응답 구성 요소
- reflection : 상황 요약/감정 반사
- anchor : 오늘 붙잡을 핵심 한 줄
-… See the full description on the dataset page: https://huggingface.co/datasets/nowsika/NOBLE_GoldenSet-KR_Verified-v3.2.1.
