datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multilingal-sakalt-dataマルチリンガルデータセットです。mitライセンスです。
JMID
JMID: Japanese Medical Incident Dataset
日本語
本データセットは、公益財団法人日本医療機能評価機構の医療事故報告書に書かれている医療事故内容から、医療事故の「具体的内容」「背景・要因」「改善策」とその他の情報をまとめたものである。
使い方の例は以下に載せる。
English
This dataset is compiled from the medical incident reports published by the Japan Council for Quality Health Care. It summarizes the contents of medical incidents, including the specific details, background and contributing factors, and proposed improvements, along with other related information.
An example of how to use the… See the full description on the dataset page: https://huggingface.co/datasets/Sakaji-Lab/JMID.sql-create-context-thai
Overview
This dataset builds from sql-create-context.
@misc{b-mc2_2023_sql-create-context,
title = {sql-create-context Dataset},
author = {b-mc2},
year = {2023},
url = {https://huggingface.co/datasets/b-mc2/sql-create-context},
note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.},
}
sakksa
Superior-Reasoning-SFT-gpt-oss-120b
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/sakksa.ColiFormer-Data
ColiFormer Training and Evaluation Dataset
This dataset contains the training and evaluation data used for the ColiFormer model - a specialized codon optimization transformer fine-tuned for Escherichia coli sequences. The model achieves 6.2% better CAI (Codon Adaptation Index) scores compared to the base CodonTransformer model.
🔗 Related Resources
Model: saketh11/ColiFormer
Base Model: adibvafa/CodonTransformer
Paper: CodonTransformer: The Global Codon Optimization… See the full description on the dataset page: https://huggingface.co/datasets/saketh11/ColiFormer-Data.sakhi-asha-home-visit-conversations
Sakhi — ASHA Home-Visit Conversations (Hindi/Hinglish → Structured Forms)
Synthetic Hindi/Hinglish conversations between an Indian ASHA (Accredited Social
Health Activist) and a patient during a maternal- and child-health home visit, each
paired with a structured JSON target. Built for the Sakhi project — an offline
voice-to-form tool for ASHA workers (github.com/Tushar-9802/Sakhi).
The dataset supports two supervised tasks over the same conversations:
form_extraction — extract a… See the full description on the dataset page: https://huggingface.co/datasets/Tushar9802/sakhi-asha-home-visit-conversations.SpeakMK1_SLP_Dialogue
Dataset Card for SLP Dialogue Dataset
Dataset Description
This dataset consists of 1,000 multi-turn simulated pediatric Speech-Language Pathology (SLP) interaction dialogues. It is specifically designed to train, evaluate, or fine-tune LLMs to act as clinical speech-language therapists or to study clinical reasoning during speech therapy sessions. Each conversation includes clinical metadata (child age, speech sound disorder category, specific phone error, clinical goal… See the full description on the dataset page: https://huggingface.co/datasets/SakhrML/SpeakMK1_SLP_Dialogue.Abkhaz-chatgptこちらはchatgptに生成してもらったサンプルです。
This is a sample generated by chatgpt.
saishin-abalpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/Sakshibeniwal10/alpaca-cleaned.
