CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cadsy /cad-technical-drawings CAD Technical Drawings, Generated by Cadsy Turn a STEP model into a labeled technical drawing automatically. This sample was created with Cadsy from 3D models in the Zero-to-CAD-100k dataset. For every STEP model, Cadsy generated: one drawing using an ASME-style profile; one drawing using an ISO-style profile; and structured bounding-box labels for every retained annotation. That is 65 CAD models, 130 technical drawings and their labels, produced through one repeatable… See the full description on the dataset page: https://huggingface.co/datasets/cadsy/cad-technical-drawings.imagen<1K1 likes1.4k downloads20d agoHugging Face02ajibawa-2023 /Technical-Architectures-Large Technical Architectures Large (294k Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.tabulartext-generation100K<n<1M8 likes248 downloads2mo agoHugging Face03electroglyph /technicalthis is a very simple dataset i created as a test, it's not really useful for much now, but it did improve benchmarks for some older embedding models. i was just attempting to do a quick and dirty expansion of a model's vocabulary. it's 100% synthetic data based on lots of occupations and the tools and terms they might use in their profession. in addition to that i added some sci-fi and fantasy terms just for laughs =) textsentence-similarity100K<n<1M4 likes187 downloads9mo agoHugging Face04stindardlogic /technical-writing-sft-100k Technical Writing SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use. Motivation Technical writing is one of the most underserved capabilities in LLMs. Common model failures: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.texttext-generation100K<n<1M0 likes187 downloads2mo agoHugging Face05lianghsun /chinese-english-technical-patent-glossary Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performing Operations; Transporting) C — 化學、冶金(Chemistry; Metallurgy) D… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/chinese-english-technical-patent-glossary.texttranslation1M<n<10M2 likes106 downloads5mo agoHugging Face06trjxter /Kimi-K2.6-Technical-Reasoning-AddOn-3300x Kimi-K2.6-Technical-Reasoning-AddOn-3300x This dataset is a technical reasoning add-on dataset generated with Kimi K2.6 as the teacher model. The dataset was designed as an additional technical reasoning trace set for downstream SFT experiments, especially around math, graduate-level science, coding, and debugging/code-repair style prompts. Dataset Summary Dataset name: Kimi-K2.6-Technical-Reasoning-AddOn-3300x Teacher model: Kimi-K2.6 Backend: W&B… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Technical-Reasoning-AddOn-3300x.texttext-generation1K<n<10K1 likes58 downloads4mo agoHugging Face07meatfly /technical_documents Technical Documents: Shafts Process (v1, PNG) license: cc-by-4.0 pretty_name: Technical Documents: Shafts Process (v1, PNG) language: - en task_categories: - visual-question-answering size_categories: - 1K<n<10K annotations_creators: - machine-generated source_datasets: - original ~1000 image→text pairs (stepped shafts → machining process) Fixed 10–11 step template; right-side Z=0; tailstock rule (L/D>3 or total length > 190 mm)Images converted from SVG to PNG… See the full description on the dataset page: https://huggingface.co/datasets/meatfly/technical_documents.image1K<n<10K0 likes48 downloads1y agoHugging Face08guanvireak /khmer-nlp-technical-corpus khmer-nlp-technical-corpus — Khmer Strategic NLP Corpus Dataset Summary This dataset contains peer-grade long-form technical treatises (3,000+ words each) in the Khmer language (km / ភាសាខ្មែរ). Every article is normalized and features neural BiGRU+CRF word segmentation with Zero-Width Space (\u200B) injection to prevent token fragmentation in sub-word tokenizers. Dataset Statistics Total Documents: 3 Train Documents: 3 Total Words: 8,002 Total… See the full description on the dataset page: https://huggingface.co/datasets/guanvireak/khmer-nlp-technical-corpus.tabulartext-generationn<1K0 likes40 downloads9d agoHugging Face09CircularBalls /tt633-technical-code-assistant-v1 TT633 Technical Code Assistant v1 This dataset is built for training the fresh custom TransformerTechnology V8.3 MDL Circle-Switch-Grid model as a small technical/code assistant. Canonical training column: text. Format: Instruction: ... Input: ... Answer: ... <END> Primary sources: Plaincode CNL rows from CircularBalls/plaincode-cnl-100k. Small curated technical QA, code-generation, debugging, reasoning, and stop-discipline seed rows. Optional local pack text if provided at… See the full description on the dataset page: https://huggingface.co/datasets/CircularBalls/tt633-technical-code-assistant-v1.texttext-generation10K<n<100K0 likes33 downloads4mo agoHugging Face10dzur658 /ping-technical-assistant-small Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine tuning. How to Utilize this Dataset In theory this dataset should work properly with… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-small.texttext-generation1K<n<10K0 likes32 downloads7mo agoHugging Face11SAWithanage /en-si-translation-weblate-technical-1k En Si Translation Weblate Technical 1K Dataset Summary English-Sinhala Technical and UI Localization dataset with ~1,000 rows targeting software interfaces, technical terminology, and static variables. Engineering Pipeline Parameters Language Pair: English (en) to Sinhala (si) Total Valid Token Rows: 1000 Internal Storage Structure: Single-File data.json Upstream Source Attribution This specific sub-split was compiled and extracted from the… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-weblate-technical-1k.texttranslation1K<n<10K0 likes26 downloads4mo agoHugging Face12sabin1234 /nepali-it-technical-sft Nepali IT SFT Dataset Cleaning Pipeline Overview This repository contains the complete cleaning, validation, improvement, and Nepali-only translation pipeline for the Nepali IT SFT (Supervised Fine-Tuning) Dataset. Starting from a raw JSONL dataset of 1,200 Nepali IT Q&A entries, this pipeline produces a production-ready, pure Nepali dataset suitable for fine-tuning large language models (LLMs). About the Dataset The dataset is a Nepali-language IT… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/nepali-it-technical-sft.text1K<n<10K0 likes25 downloads2mo agoHugging Face13dzur658 /ping-technical-assistant-mediumNow 3x the size of Ping Technical Assitant Small! NOTE: A new LoRA will be trained on this data soon! Ping Technical Assistant Dataset Small This is the dataset that was used to create Ping Technical Assistant LoRA which is an agent that focuses on technical support for consumer devices. It consists of a training dataset, validation dataset, and test dataset. The dataset is ready immediately for fine tuning tasks in MLX, and follows the format laid out by the example docs for fine… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/ping-technical-assistant-medium.texttext-generation1K<n<10K0 likes24 downloads7mo agoHugging Face14autoshift /Technical-Architectures-Large Technical Architectures Large (210k+ Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 210,000 distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/autoshift/Technical-Architectures-Large.tabulartext-generation100K<n<1M0 likes20 downloads2mo agoHugging Face15ai-training-datasets /TechnicalSupport Technical Support Conversations Dataset This dataset contains 100 realistic technical support conversations between customers and support agents. Each dialogue is 11 turns long and covers common technical issues such as wifi disconnections, printer problems, laptop battery failures, software errors, and connectivity troubleshooting. It is ideal for training AI assistants, chatbots, and support models for IT helpdesks and technical support teams. Dataset Structure Each… See the full description on the dataset page: https://huggingface.co/datasets/ai-training-datasets/TechnicalSupport.textn<1K1 likes18 downloads7mo agoHugging Face16hunterbown /bell-labs-technical-archive Bell Labs Documents and Stuff This is a conservative public-release subset of the internal BELLA continued-pretraining corpus. It keeps the Bell-system technical material that survived a stricter final pass for public dataset hosting and removes records that still looked risky, off-scope, or too low-signal for a Hugging Face corpus listing. What is in the release Split Documents train 1220 validation 29 test 42 The release contains 1291 documents out… See the full description on the dataset page: https://huggingface.co/datasets/hunterbown/bell-labs-technical-archive.tabulartext-generation1K<n<10K0 likes16 downloads6mo agoHugging Face17kanimpurath /technicalstext100K<n<1M0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.