CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zxliu /ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models" text100K<n<1M3 likes162 downloads2y agoHugging Face02Unified-Language-Model-Alignment /Anthropic_HH_Golden Dataset Card for Anthropic_HH_Golden This dataset is constructed to test the ULMA technique as mentioned in the paper Unified Language Model Alignment with Demonstration and Point-wise Human Preference (under review, and an arxiv link will be provided soon). They show that replacing the positive samples in a preference dataset by high-quality demonstration data (golden data) greatly improves the performance of various alignment methods (RLHF, DPO, ULMA). In particular, the ULMA… See the full description on the dataset page: https://huggingface.co/datasets/Unified-Language-Model-Alignment/Anthropic_HH_Golden.text10K<n<100K39 likes144 downloads3y agoHugging Face03abidlabs /repro-how-much-can-language-models-memorize-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes75 downloads2mo agoHugging Face04xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K6 likes65 downloads3y agoHugging Face05WbjuSrceu /QwQ_32B_Preview_language_model_generation_not_confirmedtext100K<n<1M0 likes61 downloads2y agoHugging Face06robotamski /language_models_lab_2text1M<n<10M0 likes51 downloads6d agoHugging Face07linneripe /language_modelstext1M<n<10M0 likes41 downloads7d agoHugging Face08wilmamuller /languagemodelstext1M<n<10M0 likes33 downloads8d agoHugging Face09Clemylia /Color-modele-languagetextn<1K0 likes28 downloads10mo agoHugging Face10heitorefer /repro-how-good-is-post-hoc-watermarking-with-language-model-rephrasing-traces Agent traces Agent sessions published from a Trackio Logbook. text1K<n<10K0 likes27 downloads2mo agoHugging Face11dairafm05 /2-language-modelstext1M<n<10M0 likes27 downloads4d agoHugging Face12DigitalIntelligenceCenter-of-ICMM /Baize-TCM-Corpus-for-Large-Language-Models-V2 白泽中医药大模型语料库 版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究 📚 简介 “白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。 本语料库可广泛应用于: 中医药大语言模型的预训练与微调 智能问答系统开发 医学自然语言处理任务(如实体识别、关系抽取) 中医药知识图谱构建 🧩 数据内容 每条语料为一个标准的问答对,格式如下: { "instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.text10K<n<100K3 likes25 downloads1y agoHugging Face13fineset-io /protein-language-models-papers Protein Language Models Papers — FineSet A research-paper dataset on Protein Language Models Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Protein Language Models Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/protein-language-models-papers.tabulartext-classificationn<1K0 likes24 downloads3mo agoHugging Face14DigitalIntelligenceCenter-of-ICMM /Baize-TCM-Corpus-for-Large-Language-Models-V1 白泽中医药大模型语料库 版本:1.0语料数量:4,735 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究 📚 简介 “白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 4,735 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。 本语料库可广泛应用于: 中医药大语言模型的预训练与微调 智能问答系统开发 医学自然语言处理任务(如实体识别、关系抽取) 中医药知识图谱构建 🧩 数据内容 每条语料为一个标准的问答对,格式如下: { "instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?", "input": "", "output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V1.text1K<n<10K1 likes23 downloads1y agoHugging Face15LunaNguyenX /Pretrain_language_model-1BL3-competesmoe11tabularn<1K0 likes4 downloads1y agoHugging Face16LunaNguyenX /Pretrain_language_model-1BL3-competesmoe10gatedtabularn<1K0 likes1 downloads1y agoHugging Face17LunaNguyenX /Pretrain_language_model-1BL3-competesmoe12gatedtabularn<1K0 likes1 downloads1y agoHugging Face18LunaNguyenX /Pretrain_language_model-1BL3-competesmoe4gatedtabularn<1K0 likes1 downloads1y agoHugging Face19LunaNguyenX /Pretrain_language_model-1BL3-competesmoe9gatedtabularn<1K0 likes1 downloads1y agoHugging Face20beldua /indonesian-language-model-lite-atabularn<1K0 likes1 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.