datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
guidelinesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/epfl-llm/guidelines
🎉 NEW DROP 🎉 PubMed Guidelines
We just added 1627 clinical guidelines found in PubMed and PubMed Central to the dataset on December 23rd, 2023. Merry Christmas!
Clinical Guidelines
The Clinical Guidelines corpus is a new dataset of 47K clinical practice guidelines from 17 high-quality online… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/guidelines.augmented-clinical-notesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes
Augmented Clinical Notes
The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources:
Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/augmented-clinical-notes.support-json-ru
Support-JSON-RU
Synthetic Russian SaaS support data for policy-conditioned JSON decisions and draft replies. The task supplies customer text, company policies, sourced facts and available capabilities; the model predicts a nine-field decision rather than memorizing a single company's policy.
Русский SaaS-support: обращение + правила + факты → категория, приоритет, настроение, действие, черновик ответа и эскалация.
Model · Dataset files · License
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/A11Sunday/support-json-ru.Solar-Open2-120B-A15B-REAM-148E-Healing-Mix
Solar-Open2-120B-A15B-REAM-148E Healing Mix
Baekpica/Solar-Open2-120B-A15B-REAM-148E-BF16 의 REAM 병합 손상을 복구하기 위한
비공개 결정론적 힐링 믹스입니다. 모든 시퀀스는 Solar Open 2 chat template로
사전 렌더·토크나이즈되어 있고 assistant 턴에만 loss 마스크가 열려 있습니다.
왜 만들었나
계보는 upstage/Solar-Open2-250B → REAP(184E) → REAM(148E) 입니다.
REAM 이후 고정 프롬프트 A/B 관찰에서 다음이 확인됐습니다.
한국어 응답에서 반복 붕괴 (예: 동일 3-gram이 출력의 88.6% 차지)
지시 준수·포맷 이탈, typo
도메인·언어에 따라 편차가 큰 열화 (일부 프롬프트는 정상)
손상은 라우터(184→148 centroid slice)와 병합된 expert 가중치에… See the full description on the dataset page: https://huggingface.co/datasets/Baekpica/Solar-Open2-120B-A15B-REAM-148E-Healing-Mix.China-K12-STEM-10K-CoT-Reasoning
K12-STEM-CoT-Chinese
1.54M Chinese K12 STEM problems with chain-of-thought solutions, 48% with diagrams.
The largest structured Chinese math/physics/chemistry reasoning dataset.
This is a curated sample (10,000 problems) of the full 1.54M dataset available via API.
Full Dataset Access
Access the full 1,540,000+ problems via API →
This Sample
Full API
Total problems
10,025
1,540,000+
With CoT solutions
10,025
1,490,000+
With diagrams
6,093
740… See the full description on the dataset page: https://huggingface.co/datasets/a13905873166/China-K12-STEM-10K-CoT-Reasoning.OMB-Circular-A11-Section-120-Apportionment-Process
Dataset Description
The OMB Circular A-11 Section 120 Apportionment Process Question Answering Dataset is a document-grounded collection of 150 question-and-answer records concerning the federal apportionment process administered by the Office of Management and Budget.
The dataset was developed from Section 120, “Apportionment Process,” of OMB Circular No. A-11, Preparation, Submission, and Execution of the Budget. Section 120 is part of the Circular’s budget-execution… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A11-Section-120-Apportionment-Process.Asclepius-Synthetic-Clinical-NotesThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-a1/Asclepius-Synthetic-Clinical-Notes.pallasbench-robust-gpu-a100
PallasBench: Robust Pallas GPU Kernel Benchmark (A100)
39/45 kernels passing on NVIDIA A100 80GB -- the first GPU-focused evaluation of JAX Pallas kernels.
What is this?
PallasBench is a suite of 45 JAX Pallas kernels across 3 difficulty levels. The original kernels were designed for TPU and failed on GPU because Pallas compiles to Triton on NVIDIA hardware, which has strict block size limits that TPU's Mosaic compiler does not.
We fixed all 45 kernels for GPU… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/pallasbench-robust-gpu-a100.qwen3.5-397b-a17b-218xTrace of Qwen3.5 397B A17B LLM.
Data count (Total: 218):
English - 108
Russian - 110
Data is presented in {"messages":[{"role":"user", "content":"Prompt"}, {"role":"assistant", "content": "Response"}]} format and each conversation split by newline.
Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/a1b8h04i/Hinglish-Everyday-Conversations-1M.distiset-ascii-art-a1
Dataset Card for distiset-ascii-art-a1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/DominguesAddem1974/distiset-ascii-art-a1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/DominguesAddem1974/distiset-ascii-art-a1.my-distiset-4d3904d1
Dataset Card for my-distiset-4d3904d1
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/A1berto0/my-distiset-4d3904d1/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/A1berto0/my-distiset-4d3904d1.lfm2.5-8b-a1b-506xTrace of LFM2.5 8B A1B LLM made by LiquidAI.
Data is presented in ChatML format and each conversation split by newline. Ready to be used for fine-tuning.
Example:
{"messages":[{"role":"user", "content":"Hello!"}, {"role":"assistant", "content":"Hello!"}]}
Brought to you by sapbot from Romarchive
