CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cadsy /cad-technical-drawings CAD Technical Drawings, Generated by Cadsy Turn a STEP model into a labeled technical drawing automatically. This sample was created with Cadsy from 3D models in the Zero-to-CAD-100k dataset. For every STEP model, Cadsy generated: one drawing using an ASME-style profile; one drawing using an ISO-style profile; and structured bounding-box labels for every retained annotation. That is 65 CAD models, 130 technical drawings and their labels, produced through one repeatable… See the full description on the dataset page: https://huggingface.co/datasets/cadsy/cad-technical-drawings.imagen<1K1 likes1.4k downloads19d agoHugging Face02gridfm /reproducibility-datakit-technical-reportSampled parquet for the gridfm-datakit technical report diversity plots. How to reproduce: scripts/datakit_report/README.md on branch genco-paper-repro. tabular100M<n<1B0 likes1.1k downloads17d agoHugging Face03MaxwellmuF /Soofi-sft-non-technical-weightedtext1M<n<10M0 likes600 downloads6mo agoHugging Face04MaxwellmuF /Soofi-sft-technical-weightedtext1M<n<10M1 likes591 downloads6mo agoHugging Face05Infatoshi /youtube_technical_v0 youtube_technical_v0 Canonical Phonon ASR dataset. Audio is embedded in Parquet as a Hugging Face-compatible audio struct with bytes and path. Hidden eval rows must not be used for training or synthetic prompt generation. Hub Dataset id: Infatoshi/youtube_technical_v0 (public) Companion labels/manifests/attribution: Infatoshi/phonon-youtube-technical Rows: see summary.json (~118k) Audio is embedded in Parquet (audio struct: bytes, path) Split… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/youtube_technical_v0.tabular100K<n<1M0 likes406 downloads2mo agoHugging Face06ajibawa-2023 /Technical-Architectures-Large Technical Architectures Large (294k Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.tabulartext-generation100K<n<1M8 likes284 downloads2mo agoHugging Face07abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes216 downloads1y agoHugging Face08Weidows /baguwen-technical-qa 八股文技术面试题库 · Baguwen Technical Interview QA 一套面向中文技术八股/面试的问答数据集,收集并整理自三个公开 GitHub 仓库,统一为 (问题, 答案) 结构化格式,适合检索增强(RAG)、SFT 微调、面试题库等场景。 数据来源(已标注) 来源 仓库 类型 处理方式 源头更新时间 状态 bestJavaer crisxuan/bestJavaer 手写 Markdown 笔记(168 篇) 直接解析原始 Markdown(## 问题 + 答案),按篇拆分,并剔除作者主观/营销干扰文本 2026-07-28 ✅ 已收录(631 条) learning_mind_map 0voice/learning_mind_map 扫描版思维导图 PDF(74 个,已成功转换 49 个) PyMuPDF 渲染每页为图片 → qwen3.8-27b 视觉 OCR → 层级 Markdown → LLM 改写为问答对 2024-05-20 🟡 部分收录(198 条;剩余… See the full description on the dataset page: https://huggingface.co/datasets/Weidows/baguwen-technical-qa.textquestion-answeringn<1K1 likes203 downloads14d agoHugging Face09QTE-Technologies /industrial-technical-archive 🚀 Latest Updates (July, 2026) Version: v07.2026 (Verified) Status: Integrated with 1,000,000+ records. New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv. QTE Technologies: Industrial & Scientific Knowledge Base Wikidata Entity: Q138411149 IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq Official Neural Hub: qtetech.github.io This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.image10K<n<100K0 likes196 downloads2mo agoHugging Face10stindardlogic /technical-writing-sft-100k Technical Writing SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use. Motivation Technical writing is one of the most underserved capabilities in LLMs. Common model failures: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.texttext-generation100K<n<1M0 likes184 downloads2mo agoHugging Face11electroglyph /technicalthis is a very simple dataset i created as a test, it's not really useful for much now, but it did improve benchmarks for some older embedding models. i was just attempting to do a quick and dirty expansion of a model's vocabulary. it's 100% synthetic data based on lots of occupations and the tools and terms they might use in their profession. in addition to that i added some sci-fi and fantasy terms just for laughs =) textsentence-similarity100K<n<1M4 likes183 downloads9mo agoHugging Face12Shekswess /technical-manuals Description Topic: Technical Manuals Domains: Engineering, Information Technology, Product Documentation Number of Entries: 1,000 Dataset Type: Raw Dataset Model Used: bedrock/us.meta.llama4-maverick-17b-instruct-v1:0 Language: English texttext-generation1K<n<10K4 likes145 downloads1y agoHugging Face13Infatoshi /phonon-youtube-technical phonon-youtube-technical This dataset contains Phonon-authored manifests, labels, term/context indexes, attribution, and local rebuild scripts for technical YouTube speech data. It intentionally contains no audio files. Contents manifests/: segment timing, source URL, video ID, YouTube-reported license, and audio SHA-256 references. labels/: Phonon-authored label queues and teacher metadata. term_context_indexes/: technical term/context indexes used for analysis… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/phonon-youtube-technical.automatic-speech-recognition0 likes111 downloads3mo agoHugging Face14sanjeevafk /depthapi_technical_corpus DepthAPI Technical Corpus Overview The DepthAPI Technical Corpus is a curated, high-quality retrieval corpus designed for modern RAG (Retrieval-Augmented Generation) systems. It features clean, aggressively normalized technical documentation, code snippets, engineering post-mortems, and system design literature. This dataset was explicitly built to serve as the local ground-truth for the DepthAPI project. About the DepthAPI Project DepthAPI is an… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/depthapi_technical_corpus.text100K<n<1M0 likes106 downloads4mo agoHugging Face15lianghsun /chinese-english-technical-patent-glossary Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performing Operations; Transporting) C — 化學、冶金(Chemistry; Metallurgy) D… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/chinese-english-technical-patent-glossary.texttranslation1M<n<10M2 likes104 downloads5mo agoHugging Face16devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes102 downloads7mo agoHugging Face17sangamdas /Apple-Siri-Europe-DMA-Interoperability-Technical-Solution Device-Side Execution-Finality Governance for On-Device AI Agents App-Scoped Fractional Capabilities, Protected Validation, and Asymmetric Operating-System Trust Overview This repository contains technical material describing a device-side execution-finality architecture for artificial-intelligence agents, on-device assistants, third-party assistants, applications, automation systems, browser-use agents, computer-use agents, application extensions… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Apple-Siri-Europe-DMA-Interoperability-Technical-Solution.documentn<1K0 likes99 downloads1mo agoHugging Face18geethanjali-cv /stock-technical-indicators Stock Technical Indicators Dataset Historical technical indicators dataset used to train directional stock movement classifiers. Features RSI: Relative Strength Index SMA_20 / EMA_50: Simple and Exponential Moving Averages MACD: Moving Average Convergence Divergence Target: Directional label (1 = Bullish, 0 = Bearish) tabulartabular-classificationn<1K1 likes92 downloads16d agoHugging Face19deerfieldgreen /stk-technical-indicators-1hourtabular10K<n<100K0 likes63 downloads2y agoHugging Face20techcwbldr /tinker-technical-training-reportsimagen<1K1 likes62 downloads4mo agoHugging Face21trjxter /Kimi-K2.6-Technical-Reasoning-AddOn-3300x Kimi-K2.6-Technical-Reasoning-AddOn-3300x This dataset is a technical reasoning add-on dataset generated with Kimi K2.6 as the teacher model. The dataset was designed as an additional technical reasoning trace set for downstream SFT experiments, especially around math, graduate-level science, coding, and debugging/code-repair style prompts. Dataset Summary Dataset name: Kimi-K2.6-Technical-Reasoning-AddOn-3300x Teacher model: Kimi-K2.6 Backend: W&B… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Technical-Reasoning-AddOn-3300x.texttext-generation1K<n<10K1 likes57 downloads4mo agoHugging Face22sangamdas /Apple-vs-the-EU-DMA-A-Technical-Way-to-Let-Third-Party-AI-In-Without-Giving-Up-Control Third Part AI Assistant Tool Interoperability Without Sacrificing Privacy or Security Execution-Finality Governance for AI Interoperability Internet-Draft: draft-das-execution-finality-ai-interoperability-01 Author: Sangam Das — Independent Inventor, Balasore, Odisha, India Draft date: 29 August 2026 Document category: Informational Internet-Draft Keywords: execution finality · AI interoperability · Digital Markets Act · non-bearer authority · Finality Sink Core principle:… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Apple-vs-the-EU-DMA-A-Technical-Way-to-Let-Third-Party-AI-In-Without-Giving-Up-Control.documentn<1K0 likes55 downloads23d agoHugging Face23sangamdas /Technical-Solution-to-Protect-Children-from-Adult-Contents-Digital-Services-Act📄 Extended Technical Publication For the detailed technical publication on the child-safety and rendering-finality architecture: DOI: https://doi.org/10.5281/zenodo.22126701 Read the full child-safety technical publication on Zenodo Age Verification Is Not Enough A Rendering-Finality Architecture to Stop 18+ and Age-Restricted Content from Reaching Children — Relevance to the EU Digital Services Act (DSA) Author / Creator: Sangam Das Role: Independent Inventor Date: 27 August 2026 Status:… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Technical-Solution-to-Protect-Children-from-Adult-Contents-Digital-Services-Act.0 likes51 downloads24d agoHugging Face24meatfly /technical_documents Technical Documents: Shafts Process (v1, PNG) license: cc-by-4.0 pretty_name: Technical Documents: Shafts Process (v1, PNG) language: - en task_categories: - visual-question-answering size_categories: - 1K<n<10K annotations_creators: - machine-generated source_datasets: - original ~1000 image→text pairs (stepped shafts → machining process) Fixed 10–11 step template; right-side Z=0; tailstock rule (L/D>3 or total length > 190 mm)Images converted from SVG to PNG… See the full description on the dataset page: https://huggingface.co/datasets/meatfly/technical_documents.image1K<n<10K0 likes48 downloads1y agoHugging Face25nirav60614 /technical-docs-qa-validated Technical Documentation Q&A - Validated This is a validated version of nirav60614/technical-docs-qa with quality scores and filtering. Validation Summary Total Pairs: 261,077 (100%) Valid Pairs: 248,096 (95.0%) Average Quality Score: 0.867/1.0 Validation Method: LLM-based (llama3.2:latest via Ollama) GPU: NVIDIA RTX 5090 Processing Time: ~28 hours Validated: 2025-11-05 Quality Distribution Quality Level Score Range Count Percentage Excellent ≥ 0.9… See the full description on the dataset page: https://huggingface.co/datasets/nirav60614/technical-docs-qa-validated.question-answering100K<n<1M0 likes42 downloads11mo agoHugging Face26guanvireak /khmer-nlp-technical-corpus khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus Dataset Summary This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers. Dataset Statistics Total Documents: 3 Train Documents: 3 Total Words: 8,002 Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.tabulartext-generationn<1K0 likes40 downloads8d agoHugging Face27SIA86 /TechnicalSupportCallstexttext-classification1K<n<10K0 likes39 downloads3y agoHugging Face28Digsm003 /cpp-security-technical-debt C/C++ Security × Technical-Debt Dataset Pilot release of 150 C function-level (vulnerable, fixed) pairs, each labeled on two axes — security mechanism and technical-debt type — and joined by an explicit debt → vulnerability link. This is not another CWE dump. Vuln corpora such as PrimeVul give a CWE but strip the commit / issue / test context that technical debt lives in. SATD corpora label debt but not the security consequence. This dataset keeps both, mined from real… See the full description on the dataset page: https://huggingface.co/datasets/Digsm003/cpp-security-technical-debt.text-classificationn<1K0 likes39 downloads6d agoHugging Face29technicalheist /cricket-alpaca Cricket Match Alpaca Dataset This dataset contains cricket match information formatted for instruction-tuning of Large Language Models (LLM) in Alpaca format. Dataset Splits Split Matches Entries Percentage Train 44,616 5,353,920 80% Valid 5,577 669,240 10% Test 5,578 669,360 10% Split Method: Match-level split (all 120 questions for a match go to the same split) Random Seed: 42 No Data Leakage: Matches are not shared across splits Data… See the full description on the dataset page: https://huggingface.co/datasets/technicalheist/cricket-alpaca.text-generation100M<n<1B0 likes38 downloads5mo agoHugging Face30dzur658 /ping-technical-assistant-small Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine tuning. How to Utilize this Dataset In theory this dataset should work properly with… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-small.texttext-generation1K<n<10K0 likes37 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.