datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VAB-vulnerability-analysis-benchmark
FBE and VAB
Two small benchmarks for security code analysis. Both grade without an LLM judge, so runs are cheap
and repeatable.
FBE (find-the-bug)
14 code snippets, each with one planted vulnerability. Ask the model to analyze the code, then check
whether it actually found the flaw.
Grading uses concept groups: the answer has to contain at least one synonym from every required group.
Four numbers come out:
found, did it identify the real vulnerability (this is… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/VAB-vulnerability-analysis-benchmark.cve-analysis
CVE & Vulnerability Analysis Dataset
A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass.
Overview
This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.Circuit-Analysis-Reasoning-Sample
⚡ EngineeringWays Data Lab: Circuit Analysis Reasoning Dataset (Free Sample)
This is a free 50-item sample of the EngineeringWays Circuit Analysis Reasoning Dataset. It is designed specifically for fine-tuning Large Language Models (LLMs) in advanced STEM problem-solving, featuring strict Chain-of-Thought (CoT) reasoning.
Want the complete, deduplicated 592-item master dataset? 👉 Get the LoRA-Ready Master File on Payhip
🚀 Dataset Overview
Most math and physics… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringWays/Circuit-Analysis-Reasoning-Sample.medqa-phi4-failure-analysisThis dataset contains a comprehensive log of reasoning and answers generated by microsoft/Phi-4-mini-instruct, evaluated on medalpaca/medical_meadow_medqa (USMLE) dataset.
This dataset represents instances where model got the answer right as well as wrong. All examples includes reasoning. The inference was performed locally on Macbook (M-series) using the MLX-LM framework (The model parameters were: temp: 0.3, max_tokens: 300).
investment_analysis
코스피 상장 기업 공시정보 기반 투자 리포트 데이터셋
이 데이터셋은 국내 코스피 상장 기업의 공시정보를 바탕으로, 투자 전문가들이 활용할 수 있는 심층적 분석과 투자 전략 제안을 목표로 제작되었습니다. 특히, 이 데이터셋은 GPT 파인튜닝에 최적화된 구조로 설계되어 있어, 다양한 역할(role)을 포함한 메시지 기반의 대화 형식으로 구성되어 있습니다.
데이터셋 구조
데이터셋은 JSONL 포맷으로 제공되며, 각 항목은 GPT 파인튜닝에 최적화된 메시지 형식을 따릅니다. 주요 구성은 다음과 같습니다:
messages: 메시지 배열 형태로 구성되어 있으며, 각 메시지는 아래와 같은 역할을 가집니다.
system: 모델의 역할과 행동 지침을 정의합니다.예시: "당신은 기업 재무 및 투자 분석 전문가입니다. 참고 컨텍스트를 기반으로 사용자 질문에 대해 정확하고 논리적으로 답변하세요."
user: 사용자의 질문과 컨텍스트(예시 데이터, 재무제표… See the full description on the dataset page: https://huggingface.co/datasets/MLOpsEngineer/investment_analysis.pt-it-analysis
