datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
QA_TaiwanEdoctor
Dataset Card for QA_TaiwanEdoctor
QA_TaiwanEdoctor 是一個以台灣線上醫療諮詢平台為來源之繁體中文醫療問答資料集,合計 178,126 筆,時間跨度 2000–2024 年。每筆包含使用者之健康提問(ask)與醫師之回覆(ans),並附上問題標題、提問日期、瀏覽次數與(可選之)評分,可作為繁中醫療 LLM 之 CPT / SFT 訓練素材。
Dataset Details
Dataset Description
繁體中文醫療問答語料長期稀缺,使得繁中 LLM 在健康諮詢場景之表現受限。本資料集整理自台灣公開線上醫療問答平台,內容涵蓋內科、外科、眼科、皮膚科、兒科、婦產科、精神科、牙科等各科別,由執業醫師回覆使用者之健康問題。
資料以 312 個 JSON 檔案儲存,每檔包含數百至上千筆 QA。每筆欄位:
title:問題標題(含編號);
ask:使用者之健康提問(自由文字);
ans:醫師之回覆;
qtime:提問日期(格式 YYYY/MM/DD);… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/QA_TaiwanEdoctor.Apertus-v1.5-QAT-10K
mlx-community/Apertus-v1.5-QAT-10K
This is a 2000 sample subset of the chosen pairs inside swiss-ai/Apertus_v1p5_Preference_Data for MLX-LM-LoRA and MLX-LoRA-Studio and the Quantization Aware Trained Appertus models.
QAT-SFT-Demo
QAT-SFT-Demo
Created with MLX-LoRA-Studio · Created with MLX LoRA Studio
Overview
Repository: Goekdeniz-Guelmez/QAT-SFT-Demo
Asset type: synthetic dataset
Created at: 2026-06-19 15:38:34 UTC
Synthetic data type: SFT
Generator model: Qwen3.5-0.8B-bf16
Samples: 100
Estimated tokens: ~17,975
This repository was prepared by MLX LoRA Studio from local training outputs.
Dataset Details
Generation type: SFT
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Goekdeniz-Guelmez/QAT-SFT-Demo.qatestset_demo
Dataset Card for GTimothee/qatestset_demo
This dataset was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems. Giskard helps identify performance, bias, and security issues in AI applications, supporting both LLM-based systems like RAG agents and traditional machine learning models for tabular data.
This dataset is a QA (Question/Answer) dataset, containing 1 pairs.
Usage
You can load this dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/qatestset_demo.qatestset_demo_gemini
Dataset Card for GTimothee/qatestset_demo_gemini
This dataset was created using the giskard library, an open-source Python framework designed to evaluate and test AI systems. Giskard helps identify performance, bias, and security issues in AI applications, supporting both LLM-based systems like RAG agents and traditional machine learning models for tabular data.
This dataset is a QA (Question/Answer) dataset, containing 1 pairs.
Usage
You can load this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/GTimothee/qatestset_demo_gemini.
