CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.7k downloads1mo agoHugging Face02cadsy /cad-technical-drawings CAD Technical Drawings, Generated by Cadsy Turn a STEP model into a labeled technical drawing automatically. This sample was created with Cadsy from 3D models in the Zero-to-CAD-100k dataset. For every STEP model, Cadsy generated: one drawing using an ASME-style profile; one drawing using an ISO-style profile; and structured bounding-box labels for every retained annotation. That is 65 CAD models, 130 technical drawings and their labels, produced through one repeatable… See the full description on the dataset page: https://huggingface.co/datasets/cadsy/cad-technical-drawings.imagen<1K1 likes1.4k downloads20d agoHugging Face03lerobot-raw /io_ai_tech_rawn<1K0 likes748 downloads2y agoHugging Face04Sachin21112004 /news-tech-datasettext10K<n<100K1 likes708 downloads57m agoHugging Face05nvidia /TechQA-RAG-Eval Dataset Description: TechQA-RAG-Eval is a reduced version of the original TechQA (IBM’s GitHub Page, HuggingFace) dataset specifically for evaluating Retrieval-Augmented Generation (RAG) systems. The dataset consists of technical support questions and their answers, sourced from real IBM developer forums where acceptable answers included links to reference technical documentation. This dataset is ready for commercial/non-commercial use. Dataset Owner(s): NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/TechQA-RAG-Eval.textn<1K10 likes593 downloads1y agoHugging Face06gussieIsASuccessfulWarlock /information_technology_instruct_mcq_2481textn<1K2 likes365 downloads2y agoHugging Face07saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes347 downloads2y agoHugging Face08TechxGenus /LeetCode-Contest LeetCode Contest Benchmark A new benchmark for evaluating Code LLMs proposed by DeepSeek-Coder, which consists of the latest algorithm problems of different difficulties. Usage git clone https://github.com/deepseek-ai/DeepSeek-Coder.git cd Evaluation/LeetCode # Set the model or path here MODEL="deepseek-ai/deepseek-coder-7b-instruct" python vllm_inference.py --model_name_or_path $MODEL --saved_path output/20240121-Jul.deepseek-coder-7b-instruct.jsonl python… See the full description on the dataset page: https://huggingface.co/datasets/TechxGenus/LeetCode-Contest.texttext-generationn<1K4 likes318 downloads2y agoHugging Face09ajibawa-2023 /Technical-Architectures-Large Technical Architectures Large (294k Samples) Overview Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8. Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.tabulartext-generation100K<n<1M8 likes284 downloads2mo agoHugging Face10elenagroundwork /groundwork-tech-2026 Groundwork Tech 2026 Open dataset for Groundwork tech pillar — 25 articles. Source: https://gworky.com/tech See data.json for records. texttext-generationn<1K0 likes246 downloads2d agoHugging Face11Techta /backend-code-generator-dataset Backend Code Generation Dataset Dataset Description This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages. Dataset Summary The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.texttext-generationn<1K1 likes219 downloads1y agoHugging Face12sarus-tech /spider_12Samples from Spider 1 and Spider 2 for SQLite. To have DBs locally for Spider 1 refer to the Getting Started of the official website. All queries here are tested in the databases in test_database. To have DBs locally for Spider 2 refer to the Quickstart of the github page (the first point is enough) text10K<n<100K0 likes204 downloads2y agoHugging Face13sarus-tech /medical_v3text100K<n<1M0 likes197 downloads2y agoHugging Face14techfren /agentbattler-bench AgentBattler Mini Ledger V5 Immutable evidence for 15/15 accepted Mini Ledger V5 runs across 3 harness × model conditions. What is here snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/site/terminal-campaign.json: compact website and analysis input. snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/campaign.json: source-revision-preserving campaign index with host paths removed. snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/runs/:… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench.tabularn<1K2 likes188 downloads2mo agoHugging Face15stindardlogic /technical-writing-sft-100k Technical Writing SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use. Motivation Technical writing is one of the most underserved capabilities in LLMs. Common model failures: Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.texttext-generation100K<n<1M0 likes184 downloads2mo agoHugging Face16electroglyph /technicalthis is a very simple dataset i created as a test, it's not really useful for much now, but it did improve benchmarks for some older embedding models. i was just attempting to do a quick and dirty expansion of a model's vocabulary. it's 100% synthetic data based on lots of occupations and the tools and terms they might use in their profession. in addition to that i added some sci-fi and fantasy terms just for laughs =) textsentence-similarity100K<n<1M4 likes183 downloads9mo agoHugging Face17Boxoffice1280 /Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques BoxOffice Verified Seeds This dataset contains the released BoxOffice seed datasets used in the benchmark pipeline described in the accompanying paper. The release includes ten verified seeds: 7 11 13 17 19 23 29 31 47 73 For each seed, we provide: a full JSONL file containing warmup rows plus evaluation rows an eval JSONL file containing only the evaluation rows a manifest JSON file a validation JSON file with directional warmup counts Layout viewer/ normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.tabulartext-generation1K<n<10K3 likes160 downloads5mo agoHugging Face18BAAI /IndustryInstruction_Technology-Research IndustryInstruction: Technology & Research This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.tabularquestion-answering100K<n<1M0 likes155 downloads1mo agoHugging Face19danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes142 downloads10mo agoHugging Face20Phase-Technologies /forge-3b-dpo-data FORGE-3B DPO Preference Data Tokenized (prompt, chosen, rejected) preference triples for DPO post-training of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2. This is data preparation output only — no model was trained to produce this. Stats Total pairs: 0 (paper target: ~200,000) Domains: 0/4 Context length: 4096 tokens (paper Appendix A.2, DPO block) Format: unpacked — one (prompt, chosen, rejected) triple per training example Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.texttext-generation100K<n<1M0 likes138 downloads3mo agoHugging Face21TechxGenus /Typst-Train Typst-Train [🤖Models] | [🛠️Code] | [📊Data] | Dataset used to train Typst-Coder, includes: 18.6K Typst texts 2.5K Markdown texts containing Typst-related content texttext-generation10K<n<100K9 likes118 downloads2y agoHugging Face22TechxGenus /deepseek_r1_code_1ktext1K<n<10K20 likes108 downloads2y agoHugging Face23lianghsun /chinese-english-technical-patent-glossary Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performing Operations; Transporting) C — 化學、冶金(Chemistry; Metallurgy) D… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/chinese-english-technical-patent-glossary.texttranslation1M<n<10M2 likes104 downloads5mo agoHugging Face24adab-tech /murya-hausa-en-lexicon-robinson1914 Robinson Hausa–English Lexicon (1914) 20,628 English→Hausa word/phrase pairs, parsed from Charles Henry Robinson's Dictionary of the Hausa Language, Volume II (English–Hausa), 3rd edition, Cambridge University Press, 1914. This is the English→Hausa training lexicon behind Murya — a Hausa-first voice assistant — used for translation fine-tuning, lexical retrieval, and Hausa embedding warm-start. Released as its own dataset so it's citable and reusable independent of the Murya… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/murya-hausa-en-lexicon-robinson1914.texttranslation10K<n<100K0 likes103 downloads1mo agoHugging Face25kurdish-tech /kurdish-grammar-eval Kurdish Grammar Minimal Pairs (BLiMP-style) — Kurmancî · Soranî A grammar-competence benchmark for Kurdish, built on the BLiMP idea: for each item, a correct sentence is paired with a corrupted version where one specific grammar rule has been deliberately broken. Score a language model by checking whether it assigns higher likelihood to the correct sentence than the corrupted one — accuracy well above 50% means the model learned the rule, not just surface fluency. Built by… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/kurdish-grammar-eval.texttext-classification1K<n<10K1 likes88 downloads1mo agoHugging Face26rojagtap /tech-qatextn<1K6 likes86 downloads3y agoHugging Face27exeyarikus /ru-tech-jobs Russian Tech Jobs Dataset Description Dataset contains job vacancy posts from one of Telegram channels focusing on IT and tech recruitment. Data Fields id: Unique identifier of the post. date: ISO timestamp of when the job was posted. text: The raw text of the post with markdown. views: The view count of the post at the time of scraping. tags: Special tags thats will be taken from text. How to use from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/exeyarikus/ru-tech-jobs.tabulartext-classification1K<n<10K0 likes84 downloads26d agoHugging Face28TechPowerB /RPRevamped-Small RPRevamped-Small-v1.0 Dataset Description RPRevamped is a synthetic dataset generated by various numbers of models. It is very diverse and is recommended if you are fine-tuning a roleplay model. This is the Small version with Medium and Tiny version currently in work. Github: RPRevamped GitHub Here are the models used in creation of this dataset: DeepSeek-V3-0324 Gemini-2.0-Flash-Thinking-Exp-01-21 DeepSeek-R1 Gemma-3-27B-it Gemma-3-12B-it Qwen2.5-VL-72B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/TechPowerB/RPRevamped-Small.texttext-generation1K<n<10K1 likes74 downloads1y agoHugging Face29tech-equity-collective /bias-correction-palestine-protocol Dataset Card for LLM Bias Correction (Palestine/Israel Context) This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel. Dataset Structure The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.texttext-generationn<1K0 likes72 downloads18d agoHugging Face30MongoDB /fake_tech_companies_market_reportstextn<1K0 likes71 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.