datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models"
repro-how-much-can-language-models-memorize-traces
Agent traces
Agent sessions published from a Trackio Logbook.
language_models_lab_2language_modelslanguagemodelsBaize-TCM-Corpus-for-Large-Language-Models-V2
白泽中医药大模型语料库
版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.Baize-TCM-Corpus-for-Large-Language-Models-V1
白泽中医药大模型语料库
版本:1.0语料数量:4,735 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 4,735 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V1.protein-language-models-papers
Protein Language Models Papers — FineSet
A research-paper dataset on Protein Language Models Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Protein Language Models Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/protein-language-models-papers.2-language-models
