datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.tw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.speech-CMMLUThis dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
Usage
python audio_evals/main.py --dataset speech-cmmlu --model MiniCPMo2_6-speech --use_model_pool --workers 2
@article{ultraevalaudio,
title={UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models},
author={Qundong Shi and Jie Zhou and Biyuan Lin and Junbo Cui and Guoyang Zeng and Yixuan Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-CMMLU.Formosa-Vision
Dataset Card for Formosa-Vision
Formosa Vision 是一份以台灣在地文化為核心的開源視覺語言資料集,從國家文化記憶庫 2.0中精選兩千餘張資料,文字描述採用 OGDL 1.0 授權、及圖片為 CC By SA(及更開放的授權條款)授權的影像,內容涵蓋景點、建築、生活場景與歷史脈絡。資料集以模型生成與人工審核並行的方式建立,透過視覺語言模型產生影像對話,再由參與者逐一檢查與修訂,確保描述的正確性、文化脈絡的一致性與語句的自然性。專案由 Twinkle AI 社群發起,結合社群協作與開放文化精神,期待成為訓練繁體中文視覺語言模型的重要基礎,幫助研究者與開發者打造能真正理解台灣文化細節的 VLM 模型。
Dataset Details
Dataset Description
Formosa Vision(又稱 台灣視覺資料集)是一個以台灣在地視覺文化為核心、集結社群力量共創的開源資料集。這項專案源自近年視覺語言模型(Vision Language Model… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/Formosa-Vision.medhallu-twins-rewritten-v2
MedHallu twins, rewritten answers
Release v2.
Twin pairs for medical hallucination detection. Each pair shares one PubMed
abstract and one question, and holds two answers: a grounded one and a
hallucinated one differing by exactly one edited span.
Both answers are generated by the same model in the same call — here
gemini-3.8-flash, prompt strict-span-v1.
Neither is taken from MedHallu or PubMedQA. This replaces
Certops/medhallu-twins-repaired-context, where the ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/Certops/medhallu-twins-rewritten-v2.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/krajavi3/Twin-2K-500.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/chadreadey/Twin-2K-500.medhallu-twins-judged
MedHallu twins, judged (claim-level faithfulness)
Splits
split
rows
train
2,998
test
500
train is drawn from pqa_artificial; test is a disjoint set of whole twin
pairs whose examples never appear in train (no twin leakage). Both are
class-balanced.
The repaired-context MedHallu twins
run through a claim-level faithfulness judge, for training a small model to
detect medical hallucinations and explain why. Each row keeps the twin
(question, answer… See the full description on the dataset page: https://huggingface.co/datasets/Certops/medhallu-twins-judged.TwinRouterBench
TwinRouterBench Static
This dataset contains the static track for TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing.
Paper: arXiv:2605.18859
Contents
data/train.parquet: Hugging Face viewer-friendly table with 970 rows. Nested fields such as messages and optional tool/function schemas are stored as JSON strings so all benchmark sources share a stable schema.
question_bank.jsonl: the original static question bank exported by… See the full description on the dataset page: https://huggingface.co/datasets/Amorph/TwinRouterBench.Twin-2K-500_edit
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Shar999/Twin-2K-500_edit.medhallu-twins-repaired-context
MedHallu twins with repaired context
Balanced medical hallucination-detection twins for training a small model to
detect hallucinated answers and explain why. Each row is a
(question, answer, context) triple labelled row_type.
Built from MedHallu, which pairs -- for the same question and source -- a
correct Ground Truth answer with a planted Hallucinated Answer. We keep
both as a twin pair, so within a pair the only difference is the hallucination.
That removes the… See the full description on the dataset page: https://huggingface.co/datasets/Certops/medhallu-twins-repaired-context.medhallu-twins-rewritten
MedHallu twins, rewritten answers
Twin pairs for medical hallucination detection. Each pair shares one PubMed
abstract and one question, and holds two answers: a grounded one and a
hallucinated one differing by exactly one edited span.
Both answers are generated by the same model in the same call. Neither is
taken from MedHallu or PubMedQA. This replaces
Certops/medhallu-twins-repaired-context,
where the ground-truth answer was a verbatim span of its own context in 99.5% of
rows… See the full description on the dataset page: https://huggingface.co/datasets/Certops/medhallu-twins-rewritten.finevision-zhtw
Dataset Card for finevisions-zhtw
Brief Summary (Preliminary)
This section describes the initial design intent of the project.
It serves as a first-edition summary and may be revised or expanded
as the dataset scope, governance, and contribution process become more clearly defined.
如何幫忙本專案
本專案正在進行中,歡迎任何夥伴加入幫忙,一起讓繁體中文的預訓練語料更多元及完善。
math-qa-ko
데이터 출처
AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 입니다.
경제 > Train > json 파일을 DataFrame 형태로 변형하여 수정 없이 업로드하였습니다.
저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다.
데이터 설명
본 데이터의 Type 은 ['양자/다자비교', '비율연산', '단서추출', '날짜추출', '가산/감산', '날짜가산/감산', '경계추출'] 로 구성되어 있습니다.
'단서추출' 데이터 예시
{'idx': 'kpf.02100351.20220202090230002',
'mediatype': '뉴스',
'medianame': '이투데이',
'category': '경제',
'source': 'https://www.etoday.co.kr/news/view/2101739',
'date': '2022-02-02',
'title': '"고객이 직접 아이디어… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-ko.F1-identity
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/F1-identity.tw-instruct-pro
Dataset Card for tw-instruct-pro
tw-instruct-pro is a Traditional Chinese (繁體中文) multi-domain instruction-following dialogue dataset.It is designed to cover both general NLP tasks and domain-specific task-oriented conversations.The dataset is intended for training and evaluating large language models in the Traditional Chinese context.
Dataset Details
Dataset Description
在 tw-instruct-pro… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-pro.math-qa-sample_addsub-ko
데이터 출처
AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다.
경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다.
Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다.
답변 생성 시 gpt-4o 를 활용했습니다.
저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다.
데이터 설명
본 데이터의 Type 은 '가산/감산' 로만 구성되어 있습니다.
데이터 예시
### context ###
2분기 순이익만 떼서 보면 증가세가 더욱 뚜렷하다. 신한금융은 9961억원, KB금융은 9911억원으로 1분기보다 각각 8.5%, 17.2% 늘었다. 하나금융은 6584억원, 우리금융은 6103억원으로 증가율은 각각 20.6%, 7.3%이다. 특히 KB금융은 분기 기준 사상 최대 실적을 올렸다.
수출 부진에 미·중… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_addsub-ko.math-qa-sample_ext-ko
데이터 출처
AI-HUB 에서 다운로드 받은 숫자연산 기계독해 데이터 를 사용해서 만든 데이터입니다.
경제 > Train > json 파일을 DataFrame 형태로 변형하여 전처리 및 답변 생성을 하였습니다.
Raw 데이터의 answer 정보를 참고하여 답변을 생성하였습니다.
답변 생성 시 gpt-4o 를 활용했습니다.
저작권에 의해 본 데이터는 외부 반출 및 타인의 acess 승낙은 불허합니다.
데이터 설명
본 데이터의 Type 은 '단서추출' 로만 구성되어 있습니다.
데이터 예시
### context ###
서울시가 민속 대명절인 추석을 맞아 내달 1일부터 20일까지 상생상회(매장), 네이버(온라인), 롯데백화점(매장)과 함께 팔도특산물로 구성된 명절 직거래장터를 진행한다고 31일 밝혔다.
팔도특산물을 구매할 수 있는 지역상생 거점공간인 '상생상회' 매장에서는 상주, 제주 등 14개 시도의 117개 농가에서 생산한 총… See the full description on the dataset page: https://huggingface.co/datasets/TwinDoc/math-qa-sample_ext-ko.
