datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.bactrainus-hotpotqa-teacher-traces
Bactrainus HotpotQA Teacher Traces
SOURCE-LINKED v1.0.0
Archived Llama 3.1 rationale and question-decomposition supervision, paired with complete SFT conversations and stable HotpotQA identities.
198,660 ROWS
4 CONFIGURATIONS
SFT MESSAGES
8B + 70B LABELS
CC BY-SA 4.0
A focused release of recovered teacher-generated supervision for multi-hop question answering. Every row contains the normalized annotation, an ordered… See the full description on the dataset page: https://huggingface.co/datasets/bactrianus/bactrainus-hotpotqa-teacher-traces.tplegacy-teachings
True Parents Legacy Teaching Archive
Digital archive of 3541 passages — sermons, speeches, prayers, and book excerpts by Sun Myung Moon (1920–2012) and Hak Ja Han Moon (1943–present), spanning 1946–2012.
Structure
Each record contains:
Field
Type
Description
id
string
URL slug (unique identifier)
title
string
Title of the sermon, speech, or passage
author
string
Speaker name
date
string
Publication date (YYYY-MM-DD)
year
string
Year only
tags… See the full description on the dataset page: https://huggingface.co/datasets/JonAuror/tplegacy-teachings.omnimcp_smartenergy_iot_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_smartenergy_iot_teaser.omnimcp_swarm_langgraph_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_swarm_langgraph_teaser.MMedBenchThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/Henrychur/MMedBench
MMedBench
💻Github Repo 🖨️arXiv Paper
The official benchmark for "Towards Building Multilingual Language Model for Medicine".
Introduction
This repo contains MMedBench, a comprehensive multilingual medical benchmark comprising 45,048 QA pairs for training and 8,518 QA pairs for testing.… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-c1/MMedBench.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.MMedBenchThis is a dataset repository made for the AISC class at Harvard Medical School. Please find the original dataset repository here: https://huggingface.co/datasets/Henrychur/MMedBench
MMedBench
💻Github Repo 🖨️arXiv Paper
The official benchmark for "Towards Building Multilingual Language Model for Medicine".
Introduction
This repo contains MMedBench, a comprehensive multilingual medical benchmark comprising 45,048 QA pairs for training and 8,518 QA pairs for testing.… See the full description on the dataset page: https://huggingface.co/datasets/aisc-team-c2/MMedBench.cpa-tax-scenarios-2026
CPA Tax Scenarios 2026
768 CPA tax impact scenarios by income, loan, filing status.
Details
Records: 768
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
Data models calculating the after-tax cost of mortgage debt across purchase, refinance, divorce buyout, and reverse mortgage scenarios. Compares itemized vs. standard deduction ($30K… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/cpa-tax-scenarios-2026.aihub_mrc_admin행정 문서 대상 기계독해 데이터
single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3146
best@32
0.5883
worst@32
0.0919
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/PS-098/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.military-base-housing-2026
Military Base Housing 2026
Complete U.S. military installation directory with VA loan analysis, BAH ranges, and unit assignments.
Details
Records: 427
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, CVLS -- NMLS #504814
Publisher: Wendy Thompson Lending Team
Coverage
Installations: 427 across all 50 states + DC, Guam, Puerto Rico
Army: 129 (forts, depots, arsenals, training centers, proving grounds)
Air… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/military-base-housing-2026.alimony-rules-by-state-2026
Alimony Rules By State 2026
Alimony/spousal support rules for 13 states with mortgage impact.
Details
Records: 13
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
State-by-state alimony duration and calculation methods mapped to mortgage qualification impact. Shows how alimony income qualifies (or disqualifies) for FHA, VA, and… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/alimony-rules-by-state-2026.military-pcs-relocation-2026
Military PCS Relocation 2026
100+ PCS relocation routes between 20 military bases.
Details
Records: 114
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
Tactical PCS relocation data mapping base-to-base moves with housing market delta analysis. Shows BAH change impact, rent-vs-buy timing for typical PCS cycles (2-3 years), and VA loan… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/military-pcs-relocation-2026.single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32
Single-turn eval — violetxi/int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
566
mean@32
0.3114
best@32
0.5795
worst@32
0.0777
pass_rate… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-int_qwen3-4b_distill_teacher_reverse_kl_lr1e-7-n32.va-funding-fees-2026
VA Funding Fees 2026
44 VA funding fee scenarios with example calculations.
Details
Records: 44
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
2026 VA fee structures layered with CPA-verified disability waiver logic to calculate long-term equity preservation for disabled veterans. Shows exact fee amounts by usage type and down payment… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/va-funding-fees-2026.bah-rates-2026
BAH Rates 2026
480 BAH rate records: 20 bases x 24 pay grade combinations.
Details
Records: 480
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
Complete BAH rate matrix cross-referenced with local housing costs to calculate 'BAH Coverage Ratio' — the percentage of PITI a service member's housing allowance covers at each base. Identifies… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/bah-rates-2026.data-teacher-seallms-v2234
Data Teacher SeaLLMs v1
Bộ dữ liệu chưng cất tri thức pháp luật Việt Nam từ mô hình SeaLLMs-v1.
Dùng cho bài toán huấn luyện SLM trong dự án thạc sĩ.
reverse-mortgage-scenarios
Reverse Mortgage Scenarios
260 HECM reverse mortgage scenarios by age, value, existing mortgage.
Details
Records: 260
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
Integrates HUD Principal Limit Factor tables with Thompson 'Sequence of Returns' modeling to show how HECM line-of-credit access hedges 401(k) depletion during market… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/reverse-mortgage-scenarios.data-teacher-seallms-v3
Data Teacher SeaLLMs v1
Bộ dữ liệu chưng cất tri thức pháp luật Việt Nam từ mô hình SeaLLMs-v1.
Dùng cho bài toán huấn luyện SLM trong dự án thạc sĩ.
