research-preview
llama-3.2-typhoon-t1-3b-research-preview-i1-GGUFtyphoon-s-thaillm-8b-instruct-research-previewopenthaigpt-thaillm-8b-instruct-v0.7.2-research-preview-ggufpagestorm-research-preview-24b-first-chapter-only-ggufpagestorm-research-preview-14b-full-booktyphoon-s-thaillm-8b-instruct-research-preview-i1-GGUFMammothModa2-Previewllama3.2-typhoon2-t1-3b-research-preview-GGUF
SR-ntsb-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of 149 rows of structured aerospace safety and engineering reasoning. It is designed to teach language models to analyze complex physical and behavioral fact patterns using standard Root Cause Analysis within the specific context of aviation accidents investigated by the National Transportation Safety Board.
The dataset is derived from real United States NTSB Aviation Accident reports. Each row provides a… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-ntsb-sft-preview-reviewed.SR-medical-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of structured clinical and forensic medical reasoning. It is designed to teach language models to analyze complex patient histories, physical presentations, and diagnostic findings using systematic differential analysis and pathophysiological synthesis within the context of peer-reviewed medical literature and case reports.
The dataset is derived from real, open-access PubMed Central (PMC) medical case reports.… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-medical-sft-preview-reviewed.liveVQA-Arxiv-Research-PreviewLiveVQA-Research-Preview
LIVEVQA: Live Visual Knowledge Seeking
Dataset Description
LIVEVQA is a benchmark dataset designed to evaluate the capabilities of Multimodal Large Language Models (MLLMs) in understanding and reasoning about live visual knowledge. Sourced from recent news articles (collected between March 14 and March 23, 2025), the dataset challenges models with questions requiring up-to-date, real-world information derived from images and associated news context.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ONE-Lab/LiveVQA-Research-Preview.typhoon-t1-3b-research-preview-data
Typhoon T1 3B Research Preview Data
Overview
This is a dataset used to train our first open reasoning model, Typhoon T1 (Research Preview): llama-3.2-typhoon-t1-3b-research-preview. It's available in Alpaca format ({instruction, input, output}), although input for all records is null. We acknowledge the owners of the original data sources. Please visit our technical blog for more details on the original data sources.
Data Splits
This dataset consists of 55… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-t1-3b-research-preview-data.SR-gao-sft-preview-reviewed
Overview
This is a Supervised Fine-Tuning preview dataset consisting of 112 rows of structured legal reasoning. It is designed to teach language models to analyze legal fact patterns using a standard legal argumentation structure (Issue/Facts, Rule, Application/Analysis, Conclusion) within the specific context of Federal Governement Bid Protest decisions.
The dataset is derived from real United States Government Accountability Office (GAO) Bid Protest decisions. Each row… See the full description on the dataset page: https://huggingface.co/datasets/Sabr-Research/SR-gao-sft-preview-reviewed.
