datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-bar-exam-hard-current-law-precedent-sft-1000
Korean Current-Law Bar Exam Hard SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 고난도 스타일 SFT 데이터 1,000문항입니다.
초기 직접 조문확인형 생성본은 실제 제14ㆍ15회 변호사시험보다 쉬워서, 이 버전은 다음 기준으로 다시 만들었습니다.
ㄱ/ㄴ/ㄷ/ㄹ 복합정오형 중심
甲/乙/丙, 검사ㆍ사법경찰관ㆍ행정청ㆍ회사ㆍ소송당사자 등이 등장하는 사례형 비중 확대
단순 근거 조문 선택형 제거
정답뿐 아니라 각 지문별 O/X 이유와 참고 법령 조문 제공
제15회 변호사시험 data/questions.csv와 높은 유사도 문항 제외
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 난도 관련… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-bar-exam-hard-current-law-precedent-sft-1000.KorfinQA
FinQA 한국어 번역본
Question, Answer 총 6252 Rows
korean-industrial-intelligence-bazaar
🇰🇷 Korea High-Value Industrial Intelligence & Supply-Chain (10,000 Sample Edition)
This dataset provides a curated 10,000-record premium showcase of South Korea's high-value industrial supply-chain, market-share, and technological intelligence.
⚡ Need the full 14,300,000+ real-time database?Query our live multi-channel B2A API Gateway directly for 0.01 USDC / USDT per query (Base L2 & Solana):Official Live API: https://husband-voltage-bass-incidents.trycloudflare.com… See the full description on the dataset page: https://huggingface.co/datasets/jkyung2/korean-industrial-intelligence-bazaar.korean-current-law-bar-exam-sft-1000
Korean Current-Law Bar Exam SFT 1000
대한민국 현행 법령을 기준으로 만든 변호사시험 선택형 스타일 SFT 데이터 1,000문항입니다.
이 데이터셋은 법무부 기출문제를 복제하지 않습니다. 기존 gyung/korean-bar-exam-moj-multiple-choice의 data/questions.csv는 난도와 과목 분포 참고 및 제15회 중복 방지 기준으로만 사용했습니다.
Files
data/questions.csv: Hugging Face preview용 메인 CSV입니다.
sft/train.jsonl: messages 형식 SFT용 JSONL입니다.
metadata/qa_report.json: 생성 수량, 과목 분포, 제15회 유사도 QA 결과입니다.
Columns
question_text: 문제와 5개 선택지
answer: 정답 번호, 1부터 5… See the full description on the dataset page: https://huggingface.co/datasets/gyung/korean-current-law-bar-exam-sft-1000.MMAD
MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
💡 This dataset is the full version of MMAD
Content:Containing both questions, images, and captions.
Questions: All questions are presented in a multiple-choice format with manual verification, including options and answers.
Images:Images are collected from the following links:
DS-MVTec
, MVTec-AD
, MVTec-LOCO
, VisA
, GoodsAD.
We retained the mask… See the full description on the dataset page: https://huggingface.co/datasets/korea2000/MMAD.KORA-Benchmark
KORA Benchmark
Resources for reproducing KORA: Adaptive Multi-Agent Orchestrated Retrieval over Knowledge Graphs — including the BioCQ benchmark dataset, entity resolution indexes, and the combined biomedical knowledge graph.
Repository Contents
Path
Description
benchmark/
BioCQ question splits (train / val / test / full)
indexes/scispacy*/
Pre-built SciSpaCy entity resolution indexes (~1 GB)
indexes/ark_bm25/
Pre-built ARK BM25 retrieval indexes… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-kgqa/KORA-Benchmark.
