CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01daaain /swebench-verified-deepseek-v4-flash-failure-analysis SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model driven by mini-swe-agent, graded with the official SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative root-cause diagnosis. Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.tabulartext-generationn<1K0 likes137 downloads3mo agoHugging Face02MK4-Research /VAB-vulnerability-analysis-benchmark FBE and VAB Two small benchmarks for security code analysis. Both grade without an LLM judge, so runs are cheap and repeatable. FBE (find-the-bug) 14 code snippets, each with one planted vulnerability. Ask the model to analyze the code, then check whether it actually found the flaw. Grading uses concept groups: the answer has to contain at least one synonym from every required group. Four numbers come out: found, did it identify the real vulnerability (this is… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/VAB-vulnerability-analysis-benchmark.textquestion-answeringn<1K0 likes55 downloads2mo agoHugging Face03massines3a /chainscope-analysis ChainScope Qwen3-8B Faithfulness Analysis Dataset This dataset contains Chain-of-Thought (CoT) faithfulness evaluation data for Qwen3-8B, including hidden state activations, labeled sentences, and evaluation results. Dataset Description We evaluated CoT faithfulness using the ChainScope methodology: Generate CoT responses for comparison questions (e.g., "Is A > B?") Generate "reversed" CoT responses for the opposite question ("Is B > A?") Compare whether the model's… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/chainscope-analysis.texttext-generation10K<n<100K0 likes49 downloads8mo agoHugging Face04mangesh-ux /logistics-cx-transcript-analysis-chatml OmniCX Logistics CX Dataset (Research Preview) Dataset Summary This dataset is designed for structured extraction of logistics and customer-experience (CX) signals from multi-turn support conversations. Each record uses ChatML-style messages with: a fixed system instruction a user transcript an assistant JSON payload matching LogisticsCXMetrics This release is a research preview and should not be treated as a production-certified benchmark. Project repository:… See the full description on the dataset page: https://huggingface.co/datasets/mangesh-ux/logistics-cx-transcript-analysis-chatml.texttext-generationn<1K0 likes47 downloads6mo agoHugging Face05sh111111111111111 /cve-analysis CVE & Vulnerability Analysis Dataset A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass. Overview This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.texttext-generation1K<n<10K1 likes45 downloads6mo agoHugging Face06stindardlogic /financial-analysis-sft-100k Financial Analysis SFT (100K) 100,000 ShareGPT conversations demonstrating expert-level financial analysis across DCF modeling, unit economics, LBO analysis, credit analysis, earnings interpretation, comparable company analysis, and financial ratio analysis. Motivation Financial AI is a critical enterprise capability — investment analysts, CFOs, startup founders, and finance teams need models that can reason through complex financial questions with the precision… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/financial-analysis-sft-100k.texttext-generation100K<n<1M1 likes45 downloads2mo agoHugging Face07EngineeringWays /Circuit-Analysis-Reasoning-Sample ⚡ EngineeringWays Data Lab: Circuit Analysis Reasoning Dataset (Free Sample) This is a free 50-item sample of the EngineeringWays Circuit Analysis Reasoning Dataset. It is designed specifically for fine-tuning Large Language Models (LLMs) in advanced STEM problem-solving, featuring strict Chain-of-Thought (CoT) reasoning. Want the complete, deduplicated 592-item master dataset? 👉 Get the LoRA-Ready Master File on Payhip 🚀 Dataset Overview Most math and physics… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringWays/Circuit-Analysis-Reasoning-Sample.texttext-generationn<1K1 likes36 downloads5mo agoHugging Face08miscusi /adaption-market-analysis-sec Market Analysis & News Instruction Dataset (SEC XBRL-grounded) Instruction-tuning data for financial analysis — fundamentals, growth and ratio arithmetic, trend and risk reading, filing navigation and comparability caveats — built from real XBRL facts, with every stated figure independently re-derived. Built for the Adaption Labs AutoScientist Challenge Part 2, Market Analysis & News track. What is in it Rows 5,068 (4,501 train / 567 eval) Task… See the full description on the dataset page: https://huggingface.co/datasets/miscusi/adaption-market-analysis-sec.texttext-generation1K<n<10K0 likes31 downloads2mo agoHugging Face09Papajams /autoscientist-market-analysis-lenitnes-dataset autoscientist-market-analysis-lenitnes-dataset The adapted dataset used to fine-tune Papajams/autoscientist-market-analysis-lenitnes for the Adaption Labs AutoScientist Challenge Part 2 (Market-Analysis & News category). Composition Total rows 27,965 Real seed rows (production DB) 1002 Unique source signals 272 Augmented rows (~19K domain + ~8K diversity) AutoScientist-augmented Seed provenance (real data): the lenitnes production platform… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/autoscientist-market-analysis-lenitnes-dataset.texttext-generation10K<n<100K0 likes30 downloads2mo agoHugging Face10twistshan /realistic-niah-count-mechanism-analysis Realistic NIAH count mechanism analysis Version 2 stores the paired geometry panel once. The default geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200 discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds 1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now one row rather than two duplicated mode rows. The common row contains the passage, gold records, slots, active needle spans, hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.tabulartext-generationn<1K0 likes29 downloads1mo agoHugging Face11kilicai /turkish-doc-summary-review-analysis-30k-jsonl ⚠️ Superseded by v2 Bu v1 dataset'te exact duplicate yoktu; ancak belge ve cevap şablonları fazla tekrar ediyordu. Güncel v2 sürümünü kullanın: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2 Generated by ML Intern This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub. Try ML Intern: https://smolagents-ml-intern.hf.space Source code:… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl.texttext-generation10K<n<100K0 likes23 downloads4mo agoHugging Face12kilicai /turkish-doc-summary-review-analysis-30k-jsonl-v2 Turkish Document Summary Review Analysis 30K JSONL v2 Tek dosya: train.jsonl. Bu v2 sürümü, v1'de görülen tekrar sorununu çözmek için yeniden üretildi: Daha fazla belge türü: proje önerisi, tutanak, denetim notu, şikâyet dosyası, politika taslağı, saha raporu, bütçe değerlendirmesi, risk kayıt formu, karar destek belgesi, olay inceleme raporu vb. Daha fazla alt görev: 30 farklı task_type. Exact duplicate + semantic template duplicate kontrolü. Cevap şablonları belgeye özel risk… See the full description on the dataset page: https://huggingface.co/datasets/kilicai/turkish-doc-summary-review-analysis-30k-jsonl-v2.texttext-generation10K<n<100K0 likes21 downloads4mo agoHugging Face13VaisakhKrishna /Emotional_Sentiment_AnalysisEmotional Sentiment Analysis Dataset for LLaMA-2 Fine-tuning (The formatted version can be directly used for fine tuning which contain only the formatted text, while the dataset.csv contain all the text, emotion, response and the formatted text) This dataset contains conversational data for training and fine-tuning language models for emotional sentiment analysis and response generation. The dataset includes user inputs, their corresponding emotional states, and tailored chatbot responses… See the full description on the dataset page: https://huggingface.co/datasets/VaisakhKrishna/Emotional_Sentiment_Analysis.texttext-classification1K<n<10K2 likes16 downloads2y agoHugging Face14Mihiret /blind-spot-analysis-gptneo Blind-Spots of GPT-Neo 1.3B Overview This dataset evaluates the EleutherAI GPT-Neo 1.3B base model by testing 10 diverse prompts in reasoning, translation, arithmetic, factual knowledge, and scientific explanation. Each prompt is evaluated against the expected correct output and blind-spot category. Model Used GPT-Neo 1.3B (Base Pre-trained Model) Source: https://huggingface.co/EleutherAI/gpt-neo-1.3B Methodology Prepare prompts targeting known… See the full description on the dataset page: https://huggingface.co/datasets/Mihiret/blind-spot-analysis-gptneo.texttext-generationn<1K0 likes15 downloads7mo agoHugging Face15ajay-pundir /e-cars-analysistexttext-generationn<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.