datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
imagenet_hard_review_data_r2Ko-StrategyQA
Ko-StrategyQA
This dataset represents a conversion of the Ko-StrategyQA dataset into the BeIR format, making it compatible for use with mteb.
The original dataset was designed for multi-hop QA, so we processed the data accordingly. First, we grouped the evidence documents tagged by annotators into sets, and excluded unit questions containing 'no_evidence' or 'operation'.
Ko-Agent-Trajectories-1.0
Ko-Agent-Trajectories-1.0
Dataset card v1.1.1 (2026-09-22). The pipeline code is now released in this repository
under pipeline/, together with the API catalogue, the scenario templates and the complete
prompt set. The card reports the completed human review study and the v1.1 artefacts
(behaviour DPO config, per-item validation scores, manifest, filter asset).
Korean edition: README.ko.md.
TL;DR
A Korean multi-turn agent ↔ tool trajectory corpus synthesized… See the full description on the dataset page: https://huggingface.co/datasets/taejoon89/Ko-Agent-Trajectories-1.0.ko-lima
Dataset Card for KoLIMA
Dataset Description
KoLIMA는 Meta에서 공개한 LIMA: Less Is More for Alignment (Zhou et al., 2023)의 학습 데이터를 한국어로 번역한 데이터셋입니다. 번역에는 DeepL API를 활용하였고, SK(주) Tech Collaborative Lab으로부터 비용을 지원받았습니다. 전체 텍스트 중에서 code block이나 수식을 나타내는 특수문자 사이의 텍스트는 원문을 유지하는 형태로 번역을 진행하였으며, train 데이터셋 1,030건과 test 데이터셋 300건으로 구성된 총 1,330건의 데이터를 활용하실 수 있습니다. 현재 동일한 번역 문장을 plain, vicuna 두 가지 포멧으로 제공합니다.
데이터셋 관련하여 문의가 있으신 경우 메일을 통해 연락주세요! 🥰
This is Korean LIMA dataset, which is… See the full description on the dataset page: https://huggingface.co/datasets/taeshahn/ko-lima.Ko-miracl
Ko-miracl
This dataset represents a conversion of the Korean (Ko) section from the miracl dataset into the BeIR format, making it compatible for use with mteb.
Ko-mrtydi
Ko-mrtydi
This dataset represents a conversion of the Korean (Ko) section from the Mr.TyDI dataset into the BeIR format, making it compatible for use with mteb.
domeggook_faqling-lang-chat-sft
Ling-Lang Chat SFT Dataset
Instruction-tuning data (chat/messages JSONL) for fine-tuning LLMs to help with the
Ling programming language — a multilingual systems
programming language with keyword aliases in 15 human languages, a Cranelift
JIT/AOT + tree-walking interpreter + bytecode VM, and a ling-* crate ecosystem.
Used to fine-tune zai-org/GLM-4-32B-0414
via QLoRA for chat.ling-lang.org.
Categories
File
Topic
01_syntax.jsonl
Core ling-lang syntax… See the full description on the dataset page: https://huggingface.co/datasets/taellinglin/ling-lang-chat-sft.test_amazonAdvisingNetworksReviewDataExtensionhanabi-zscfinal-qwen3-8b-logs
zsc_final Qwen3-8B 실행 로그 — 서버 10.45.7.134
llm_agent 저장소 zsc_final 에서 2026-09-21~22 에 이 서버가 돌린 것들의
로그다. 수치의 정본은 저장소의 zsc_final/result_loop_trial.md 와
zsc_final/result_hidden_store.md 다.
무엇이 들어 있나
폴더
무엇
logs/ev_*.log
LLM 게임 세 조건 (A 셀 · scope 셀 · C 셀), 판 3000~3059, 조건당 LLM 게임 240개
logs/store_game_s*.log
은닉 저장소 1차 실행 (조각 셋)
logs/store_game_t*.log
은닉 저장소 본 실행 (조각 여섯), LLM 게임 240개
logs/hdr_cut2.log
헤더 토막을 빼며 판독 정확도를 잰 실행
store_index/index_*of6.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/taekbae/hanabi-zscfinal-qwen3-8b-logs.hanabi-zscfinal-hidden-store-qwen3-8b
Hanabi run_v1 LLM 게임 은닉 저장소 — Qwen3-8B, A 셀(raw)
llm_agent 저장소 zsc_final 의 src/env2_hidden_store.py (git 107a5a8)
로 2026-09-22 에 서버 10.45.7.134 에서 뽑았다. 수치의 정본은 저장소의
zsc_final/result_hidden_store.md 다.
무대
Qwen3-8B. 측정 틀 run_v1 — p_v2 프롬프트 헤더, 발신은 양 자리 모두 대본
(규칙대로만 낸다), 수신만 LLM. 조건은 A 셀 (raw, 수신에 아무 문장도 안
넣는 무개입). 판 3000~3059 가 LLM 게임 판이고, 판 하나가 규약 구성
4개에서 돌아 LLM 게임 4개가 된다. 판 60 × 구성 4 = LLM 게임 240.
파일
raw_c{구성}_e{판}_s{좌석}.npz — LLM 게임 하나에 좌석 둘, 합쳐… See the full description on the dataset page: https://huggingface.co/datasets/taekbae/hanabi-zscfinal-hidden-store-qwen3-8b.CLIcK
This dataset is the same as https://huggingface.co/datasets/EunsuKim/CLIcK. This dataset has been subdivided for simplified viewing and evaluation.
CLIcK 🇰🇷🧠
Evaluation of Cultural and Linguistic Intelligence in Korean
Introduction 🎉
CLIcK (Cultural and Linguistic Intelligence in Korean) is a comprehensive dataset designed to evaluate cultural and linguistic intelligence in the context of Korean language models. In an era where diverse… See the full description on the dataset page: https://huggingface.co/datasets/taeminlee/CLIcK.korean-service-query-gating-v1
Korean Service Query Gating V1
한국어 서비스 질의 게이팅 연구를 위한 합성 데이터셋입니다.
현실의 서비스 회사에서는 질문 분해, 도메인 판별, 라우팅 같은 태스크가 매우 중요하지만, 실제 질의 로그는 개인정보와 내부 정책 문제 때문에 공개하기 어렵습니다.이 데이터셋은 그런 제약 아래에서도 재현 가능한 연구를 시작할 수 있도록 만든 공개 가능한 Korean starting benchmark입니다.
쉽게 말해 이 데이터셋은 아래 질문을 연구하기 위한 리소스입니다.
“이 질문이 우리 서비스 범위 안인가, 밖인가?”
“겉보기엔 비슷한데 실제 의미는 다른 질문을 어떻게 구분할까?”
“모델이 어떤 실패 유형에서 흔들리는가?”
Quick Summary
언어: 한국어
용도: 서비스 질의 게이팅 / 헬프데스크 질의 분류 / OOD stress evaluation
구성: core + stress_eval
성격: 합성 연구용… See the full description on the dataset page: https://huggingface.co/datasets/taeyun16/korean-service-query-gating-v1.KoAlpaca_hira_v1.1a
Dataset Card for "KoAlpaca-v1.1a"
Project Repo
Github Repo: Beomi/KoAlpaca
How to use
>>> from datasets import load_dataset
>>> ds = load_dataset("beomi/KoAlpaca-v1.1a", split="train")
>>> ds
Dataset({
features: ['instruction', 'input', 'output'],
num_rows: 21272
})
>>> ds[0]
{'instruction': '양파는 어떤 식물 부위인가요? 그리고 고구마는 뿌리인가요?',
'output': '양파는 잎이 아닌 식물의 줄기 부분입니다. 고구마는 식물의 뿌리 부분입니다. \n\n식물의 부위의 구분에 대해 궁금해하는 분이라면 분명 이 질문에 대한 답을 찾고 있을 것입니다. 양파는 잎이 아닌 줄기… See the full description on the dataset page: https://huggingface.co/datasets/Taegyuu/KoAlpaca_hira_v1.1a.KoAlpaca-v1.1a
Dataset Card for "KoAlpaca-v1.1a"
Project Repo
Github Repo: Beomi/KoAlpaca
How to use
>>> from datasets import load_dataset
>>> ds = load_dataset("beomi/KoAlpaca-v1.1a", split="train")
>>> ds
Dataset({
features: ['instruction', 'input', 'output'],
num_rows: 21155
})
>>> ds[0]
{'instruction': '양파는 어떤 식물 부위인가요? 그리고 고구마는 뿌리인가요?',
'output': '양파는 잎이 아닌 식물의 줄기 부분입니다. 고구마는 식물의 뿌리 부분입니다. \n\n식물의 부위의 구분에 대해 궁금해하는 분이라면 분명 이 질문에 대한 답을 찾고 있을 것입니다. 양파는 잎이 아닌 줄기… See the full description on the dataset page: https://huggingface.co/datasets/Taegyuu/KoAlpaca-v1.1a.imagenet_hard_review_datawikitableqa-koBlindTest-GeneratedDataDGB_ProjecticecompanyAdvisingNetworksReviewInitialRunsAdvisingNetworksReviewDatamoonboard_dataset
