datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science-on-a-sphere-prompt-completions
Dataset Card for Science On a Sphere QA Dataset
Dataset Details
Dataset Description
This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science.
This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.figment-eval-traces
Figment Eval Traces
Synthetic and de-identified evaluation traces for Figment, a prototype protocol-navigation aid for trained rural-clinic and disaster-response field responders.
These records are intended for model and harness debugging. They are not clinical data, medical advice, diagnosis, treatment instructions, or a substitute for local protocol, clinician judgment, supervisor review, or trained responder judgment.
Dataset Summary
The dataset captures… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/figment-eval-traces.TAA-ChatML-10K
TAA-ChatML-10K
A dataset of 10,438 question-answer pairs for Cyber Threat Intelligence (CTI) and Advanced Persistent Threat (APT) attribution tasks. Synthesized from 1,468 publicly available threat intelligence reports covering APT attribution, malware analysis, and threat actor TTPs. The dataset is formatted in ChatML conversation structure for fine-tuning large language models.
License
MIT
house-hacking-roi-scenarios
House Hacking ROI Scenarios
72 duplex/triplex/fourplex ROI calculations for house hackers.
Details
Records: 72
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Beau Thompson, NMLS #1615561
Publisher: Good News Lending
Thompson Alpha Logic
Comparative yield data for 2-4 unit properties, calculating the 'Tenant Offset Ratio' — the percentage of PITI covered by rental income. Shows the true cost of living for house hackers using FHA 3.5%… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/house-hacking-roi-scenarios.PaperProf-traces
PaperProf Agent Trace
Step-by-step trace of PaperProf,
an AI study buddy that turns course PDFs into interactive quiz sessions.
What's in this dataset
Each row in paperprof_trace.jsonl is one LLM call. Fields:
Field
Description
session_id
Groups steps from the same session
step
Step index within the session (1–4)
type
question_generation / answer_evaluation / mcq_generation
topic
Domain of the source chunk
input
Exact input sent to the model… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/PaperProf-traces.packetcourt-golden-cases
PacketCourt Golden Cases
A small evidence-first evaluation set for auditing front-of-pack claims against
the text printed on the same Indian packaged-food label.
Each record contains:
front-label claim text
back-label evidence text
expected claims and conservative verdicts
expected persuasion-gap concepts
expected deterministic date or whole-packet calculations
The initial set is intentionally small and hand-audited. It is a regression and
demonstration asset, not a… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/packetcourt-golden-cases.
