datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NLU-Sentiment-Analysis
SEA Sentiment Analysis
SEA Sentiment Analysis evaluates a model's ability to identify the sentiment polarity of a text. It is sampled from NusaX for Indonesian, Javanese, and Sundanese, IndicSentiment for Tamil, Wisesight Sentiment for Thai, and UIT-VSFC for Vietnamese.
Supported Tasks and Leaderboards
SEA Sentiment Analysis is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Sentiment-Analysis.Open-Router-API-Pricing-Analysis
OpenRouter API Pricing Analysis Dataset
Overview
This dataset provides a point-in-time capture of pricing and parameters for LLMs available through the OpenRouter API for inference.
Contents
Raw Data (raw/)
Contains the original data extracted from the OpenRouter API, including:
Model pricing (input/output token costs)
Model parameters and specifications
Computed fields such as output/input token price ratios
Enhanced Data (hf-enhanced/)… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Open-Router-API-Pricing-Analysis.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.Adversarial-Agent-Intent-Safety-Analysis-240K
Adversarial Agent Intent Safety Analysis 240K
Abstract
The Adversarial-Agent-Intent-Safety-Analysis-240K is a deterministically structured dataset featuring 242,454 context-rich adversarial prompts and safety evaluations. Engineered strictly for training frontier command-and-control models, guardrail classifiers, and red-teaming agents, it encourages models to parse multi-layered intention across 126 critical risk vectors.
This design trains models to decouple the surface… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Adversarial-Agent-Intent-Safety-Analysis-240K.VAB-vulnerability-analysis-benchmark
FBE and VAB
Two small benchmarks for security code analysis. Both grade without an LLM judge, so runs are cheap
and repeatable.
FBE (find-the-bug)
14 code snippets, each with one planted vulnerability. Ask the model to analyze the code, then check
whether it actually found the flaw.
Grading uses concept groups: the answer has to contain at least one synonym from every required group.
Four numbers come out:
found, did it identify the real vulnerability (this is… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/VAB-vulnerability-analysis-benchmark.chainscope-analysis
ChainScope Qwen3-8B Faithfulness Analysis Dataset
This dataset contains Chain-of-Thought (CoT) faithfulness evaluation data for Qwen3-8B, including hidden state activations, labeled sentences, and evaluation results.
Dataset Description
We evaluated CoT faithfulness using the ChainScope methodology:
Generate CoT responses for comparison questions (e.g., "Is A > B?")
Generate "reversed" CoT responses for the opposite question ("Is B > A?")
Compare whether the model's… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/chainscope-analysis.cve-analysis
CVE & Vulnerability Analysis Dataset
A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass.
Overview
This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.logistics-cx-transcript-analysis-chatml
OmniCX Logistics CX Dataset (Research Preview)
Dataset Summary
This dataset is designed for structured extraction of logistics and customer-experience (CX) signals from multi-turn support conversations.
Each record uses ChatML-style messages with:
a fixed system instruction
a user transcript
an assistant JSON payload matching LogisticsCXMetrics
This release is a research preview and should not be treated as a production-certified benchmark.
Project repository:… See the full description on the dataset page: https://huggingface.co/datasets/mangesh-ux/logistics-cx-transcript-analysis-chatml.financial-analysis-sft-100k
Financial Analysis SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level financial analysis across DCF modeling, unit economics, LBO analysis, credit analysis, earnings interpretation, comparable company analysis, and financial ratio analysis.
Motivation
Financial AI is a critical enterprise capability — investment analysts, CFOs, startup founders, and finance teams need models that can reason through complex financial questions with the precision… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/financial-analysis-sft-100k.tatar-news-analysis-multilabel
Dataset Card for Tatar News Multilabel Classification
Dataset Details
Dataset Description
The Tatar News Multilabel Classification Dataset contains 55,709 Tatar language news articles annotated with 13 distinct topic labels in a multi-label setting (each article can have multiple labels). Each entry includes the full article content, title, label indices, multi-hot label vector, number of labels, original single category, source URL, publication… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multilabel.adaption-market-analysis-sec
Market Analysis & News Instruction Dataset (SEC XBRL-grounded)
Instruction-tuning data for financial analysis — fundamentals, growth and ratio arithmetic, trend and risk reading, filing navigation and comparability caveats — built from real XBRL facts, with every stated figure independently re-derived.
Built for the Adaption Labs AutoScientist Challenge Part 2, Market Analysis & News track.
What is in it
Rows
5,068 (4,501 train / 567 eval)
Task… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-market-analysis-sec.autoscientist-market-analysis-lenitnes-dataset
autoscientist-market-analysis-lenitnes-dataset
The adapted dataset used to fine-tune
Papajams/autoscientist-market-analysis-lenitnes
for the Adaption Labs AutoScientist Challenge Part 2 (Market-Analysis &
News category).
Composition
Total rows
27,965
Real seed rows (production DB)
1002
Unique source signals
272
Augmented rows (~19K domain + ~8K diversity)
AutoScientist-augmented
Seed provenance (real data): the lenitnes production platform… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/autoscientist-market-analysis-lenitnes-dataset.bimcv-analysis
BIMCV-R 500-Series Multi-Model Analysis & Explainability Dataset
This repository contains end-to-end multi-model AI evaluation, lung cancer risk modeling, explainability maps (Grad-CAM), and automated radiology report generation for 500 randomly sampled physician-labeled chest CT series from the cyd0806/BIMCV-R dataset.
The cohort was evaluated across three state-of-the-art chest CT deep learning models:
Sybil (5-Seed Ensemble): 1-to-6 year lung cancer risk probability… See the full description on the dataset page: https://huggingface.co/datasets/chn123/bimcv-analysis.realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.Circuit-Analysis-Reasoning-Sample
⚡ EngineeringWays Data Lab: Circuit Analysis Reasoning Dataset (Free Sample)
This is a free 50-item sample of the EngineeringWays Circuit Analysis Reasoning Dataset. It is designed specifically for fine-tuning Large Language Models (LLMs) in advanced STEM problem-solving, featuring strict Chain-of-Thought (CoT) reasoning.
Want the complete, deduplicated 592-item master dataset? 👉 Get the LoRA-Ready Master File on Payhip
🚀 Dataset Overview
Most math and physics… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringWays/Circuit-Analysis-Reasoning-Sample.tatar-news-analysis-multiclass
Dataset Card for Tatar News Multiclass Classification
Dataset Details
Dataset Description
The Tatar News Multiclass Classification Dataset contains 86,963 Tatar language news articles classified into 9 distinct topic categories. Each entry includes the full article content, title, category (numeric label and text label), source URL, publication date, and content length. The dataset is specifically designed for training and evaluating multi-class… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatar-news-analysis-multiclass.turkish-doc-summary-review-analysis-30k-jsonl
⚠️ Superseded by v2
Bu v1 dataset'te exact duplicate yoktu; ancak belge ve cevap şablonları fazla tekrar ediyordu.
Güncel v2 sürümünü kullanın:
https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code:… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl.turkish-doc-summary-review-analysis-30k-jsonl-v2
Turkish Document Summary Review Analysis 30K JSONL v2
Tek dosya: train.jsonl.
Bu v2 sürümü, v1'de görülen tekrar sorununu çözmek için yeniden üretildi:
Daha fazla belge türü: proje önerisi, tutanak, denetim notu, şikâyet dosyası, politika taslağı, saha raporu, bütçe değerlendirmesi, risk kayıt formu, karar destek belgesi, olay inceleme raporu vb.
Daha fazla alt görev: 30 farklı task_type.
Exact duplicate + semantic template duplicate kontrolü.
Cevap şablonları belgeye özel risk… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2.Coq-Analysis
Coq-Analysis
Structured dataset from MathComp Analysis — MathComp-compatible classical real analysis.
Source
Repository: https://github.com/math-comp/analysis
Commit: 723425a8e25ee4d32ff8409d0294d25d4e43f9ad
Files: 126
License: other
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Analysis.Emotional_Sentiment_AnalysisEmotional Sentiment Analysis Dataset for LLaMA-2 Fine-tuning
(The formatted version can be directly used for fine tuning which contain only the formatted text, while the dataset.csv contain all the text, emotion, response and the formatted text)
This dataset contains conversational data for training and fine-tuning language models for emotional sentiment analysis and response generation. The dataset includes user inputs, their corresponding emotional states, and tailored chatbot responses… See the full description on the dataset page: https://huggingface.co/datasets/VaisakhKrishna/Emotional_Sentiment_Analysis.decimind-company-analysis-tr
DeciMind Company Analysis TR
Türkçe kurumsal analiz görevleri için hazırlanmış Gemma supervised fine-tuning
datasetidir.
Dataset özeti
Dil: Türkçe
Görev: supervised fine-tuning / kurumsal analiz
Toplam kayıt: 10869
Train: 8695
Validation: 1087
Test: 1087
Kaynaklar
turkish_finance_sft: 2397 kayıt
churn_prediction: 7002 kayıt
ibm_hr: 1470 kayıt
Kaynak datasetlerin lisansları ve yeniden dağıtım koşulları ayrıca kontrol
edilmelidir. FinQA retrieval… See the full description on the dataset page: https://huggingface.co/datasets/TalhaKa/decimind-company-analysis-tr.blind-spot-analysis-gptneo
Blind-Spots of GPT-Neo 1.3B
Overview
This dataset evaluates the EleutherAI GPT-Neo 1.3B base model by testing 10 diverse prompts in reasoning, translation, arithmetic, factual knowledge, and scientific explanation. Each prompt is evaluated against the expected correct output and blind-spot category.
Model Used
GPT-Neo 1.3B (Base Pre-trained Model)
Source: https://huggingface.co/EleutherAI/gpt-neo-1.3B
Methodology
Prepare prompts targeting known… See the full description on the dataset page: https://huggingface.co/datasets/Mihiret/blind-spot-analysis-gptneo.smolified-legal-smol-sovereign-contract-analysis2
🤏 smolified-legal-smol-sovereign-contract-analysis2
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model drago-2435/smolified-legal-smol-sovereign-contract-analysis2.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 87008c44)
Records: 420
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by drago-2435.
Generated… See the full description on the dataset page: https://huggingface.co/datasets/drago-2435/smolified-legal-smol-sovereign-contract-analysis2.mistakes-analysis
Nanbeige-4-3B-Base: Semantic Blind Spots & Mistake Analysis
Overview
This dataset is a curated collection of 10+ high-fidelity semantic failures identified during the evaluation of the Nanbeige/Nanbeige4-3B-Base model.
While the Nanbeige model is highly capable for its size, my testing revealed specific "blind spots" in logical reasoning, strict constraint adherence, and mathematical precision. This dataset serves as a benchmark for where the model currently fails… See the full description on the dataset page: https://huggingface.co/datasets/Odelolasolomon/mistakes-analysis.market-analysis-africa
Market Analysis & News — African Context Dataset
Instruction-tuning dataset covering African market analysis and business news: stock exchanges (USE, NSE, NGX, JSE), commodity markets (coffee, cocoa, gold, oil), regional trade (EAC, AfCFTA, ECOWAS), macroeconomic indicators, startup/VC ecosystems, real estate, and sector intelligence — grounded via web search, generated with gemini-2.5-flash.
Dataset Details
Rows: 187
Regions covered: Uganda, Kenya, Tanzania… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/market-analysis-africa.bio-image-analysis-qa
Dataset Card for bio-image-analysis-qa
This dataset contains questions and answers for analysing biological microscopy imaging data using python.
Dataset Details
Dataset Description
Questions and answers provided in this repository are centered around the topic, how to process imaging data using Python.
Curated by: Robert Haase
License: CC-BY 4.0
Dataset Sources and Processing
This dataset was derived from the Bio-image Analysis Notebooks which… See the full description on the dataset page: https://huggingface.co/datasets/haesleinhuepf/bio-image-analysis-qa.task823_peixian-rtgender_sentiment_analysis
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task823_peixian-rtgender_sentiment_analysis
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task823_peixian-rtgender_sentiment_analysis.smolified-legal-smol-sovereign-offline-contract-analysis
🤏 smolified-legal-smol-sovereign-offline-contract-analysis
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-legal-smol-sovereign-offline-contract-analysis.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 7d000d23)
Records: 1155
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.… See the full description on the dataset page: https://huggingface.co/datasets/smolify/smolified-legal-smol-sovereign-offline-contract-analysis.NCSS_2023_Data_Analysisturkish-sft-reasoning-task-analysis-30k
Turkish Reasoning + Task Following + Analysis 30K
Sıfırdan üretilmiş 30K ek SFT dataset'i.
reasoning / akıl yürütme: 10K
task-following / görev takibi: 10K
analysis / analiz: 10K
Audit
{
"rows": 30000,
"categories": {
"reasoning": 10000,
"task-following": 10000,
"analysis": 10000
},
"bad_generated": 0,
"duplicates_skipped": 0,
"constraint_failures": 0,
"reasoning_bad_answers": 0,
"analysis_missing_result": 0,
"safety_hits": 0
}… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-sft-reasoning-task-analysis-30k.
