openmed
OpenMed-NER-PharmaDetect-SuperClinical-434MOpenMed-PII-SuperClinical-Small-44M-v1OpenMed-NER-ChemicalDetect-ModernMed-149MOpenMed-NER-PharmaDetect-BigMed-278MOpenMed-NER-BloodCancerDetect-TinyMed-65MOpenMed-NER-DiseaseDetect-BioMed-335MOpenMed-NER-OncologyDetect-MultiMed-568MOpenMed-NER-AnatomyDetect-ElectraMed-109M
Datasets
All datasets matching “openmed”OpenMedText
OpenMedText Dataset
A comprehensive biomedical text corpus consisting of MDPI journal articles and open-source medical textbooks for language model training and research.
Dataset Summary
OpenMedText is a large-scale biomedical text dataset that includes:
Med-MDPI: 121,489 biomedical journal articles across 37 categories from MDPI journals
Med-Textbooks: 29 open-source medical textbooks covering various medical disciplines
Folder Structure
OpenMedText/
├──… See the full description on the dataset page: https://huggingface.co/datasets/ywchoi/OpenMedText.OpenMedallion
OpenMedallion Financial Dataset
Comprehensive financial dataset for quantitative research, machine learning, and trading strategy backtesting.
2,609 Parquet files · 18 categories · 1,500+ unique assets · 87+ countries · Up to 100 years of history
Quick Start
import pandas as pd
# Load BTC 1h data (5 years)
df = pd.read_parquet("data/BTC-BTCUSD_1h.parquet")
# Load Indonesian stocks
df = pd.read_parquet("data/equities/country_stocks/BBCA_1d.parquet")
# Load Gold… See the full description on the dataset page: https://huggingface.co/datasets/oyi77/OpenMedallion.Medical-Reasoning-SFT-Mega
Medical-Reasoning-SFT-Mega
The ultimate medical reasoning dataset - combining 7 state-of-the-art AI models with fair distribution deduplication. 1.79 million unique samples with 3.78 billion tokens of medical chain-of-thought reasoning.
Dataset Overview
Metric
Value
Total Samples
1,789,998 (after deduplication)
Total Tokens
~3.78 Billion
Content Tokens
~2.22 Billion
Reasoning Tokens
~1.56 Billion
Samples with Reasoning
1,789,764 (100.0%)
Unique… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Mega.synthvision-annotated-qwen
synthvision-annotated-qwen
Medical images annotated by Qwen 3.5 (397B) via Doubleword
Records: 59,476
About
First-half annotations from the SynthVision pipeline. 59,476 medical images annotated by Qwen 3.5 (397B MoE, 17B active) via Doubleword batch inference.
Each record contains a multi-turn clinical conversation (5-9 turns), a clinical narrative report, structured findings, reasoning chain, and difficulty rating.
Schema
id: str #… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-annotated-qwen.synthvision-validated-qwen-by-kimi
synthvision-validated-qwen-by-kimi
Qwen 3.5 annotations validated by Kimi K2.5 (93.1% pass rate)
Records: 55,359
About
Cross-validated subset from the SynthVision pipeline. Kimi K2.5 reviewed all 59,476 Qwen 3.5 annotations and confirmed 55,359 as consistent with the source images (93.1% pass rate).
Validation criteria: consistent == true AND confidence >= 0.7. Records that failed validation were removed — primarily cases where the annotator hallucinated findings not… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/synthvision-validated-qwen-by-kimi.Medical-Reasoning-SFT-GPT-OSS-120B
Medical-Reasoning-SFT-GPT-OSS-120B
A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work.
Dataset Statistics
Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.
