datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndustryCorpus_technology[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.cad-technical-drawings
CAD Technical Drawings, Generated by Cadsy
Turn a STEP model into a labeled technical drawing automatically.
This sample was created with Cadsy from 3D models in the
Zero-to-CAD-100k dataset.
For every STEP model, Cadsy generated:
one drawing using an ASME-style profile;
one drawing using an ISO-style profile; and
structured bounding-box labels for every retained annotation.
That is 65 CAD models, 130 technical drawings and their labels, produced
through one repeatable… See the full description on the dataset page: https://huggingface.co/datasets/cadsy/cad-technical-drawings.io_ai_tech_rawnews-tech-datasetTechQA-RAG-Eval
Dataset Description:
TechQA-RAG-Eval is a reduced version of the original TechQA (IBM’s GitHub Page, HuggingFace) dataset specifically for evaluating Retrieval-Augmented Generation (RAG) systems. The dataset consists of technical support questions and their answers, sourced from real IBM developer forums where acceptable answers included links to reference technical documentation.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/TechQA-RAG-Eval.information_technology_instruct_mcq_2481tech-docs
Technical Documentation Dataset
A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices.
Dataset Overview
This dataset includes documentation across multiple domains:
Cloud Platforms: GCP (83 docs), EKS (33 docs)
Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.LeetCode-Contest
LeetCode Contest Benchmark
A new benchmark for evaluating Code LLMs proposed by DeepSeek-Coder, which consists of the latest algorithm problems of different difficulties.
Usage
git clone https://github.com/deepseek-ai/DeepSeek-Coder.git
cd Evaluation/LeetCode
# Set the model or path here
MODEL="deepseek-ai/deepseek-coder-7b-instruct"
python vllm_inference.py --model_name_or_path $MODEL --saved_path output/20240121-Jul.deepseek-coder-7b-instruct.jsonl
python… See the full description on the dataset page: https://huggingface.co/datasets/TechxGenus/LeetCode-Contest.Technical-Architectures-Large
Technical Architectures Large (294k Samples)
Overview
Generating complex, syntactically valid diagram code from natural language requirements is a major challenge for AI models. This dataset bridges that gap by providing over 293,000+ distinct enterprise software architectures generated using two cutting-edge models: GPT-OSS-120B and Qwen3-Coder-Next-FP8.
Unlike simple "toy" examples, these architectures model realistic enterprise systems complete with client… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Technical-Architectures-Large.groundwork-tech-2026
Groundwork Tech 2026
Open dataset for Groundwork tech pillar — 25 articles.
Source: https://gworky.com/tech
See data.json for records.
backend-code-generator-dataset
Backend Code Generation Dataset
Dataset Description
This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages.
Dataset Summary
The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.spider_12Samples from Spider 1 and Spider 2 for SQLite.
To have DBs locally for Spider 1 refer to the Getting Started of the official website.
All queries here are tested in the databases in test_database.
To have DBs locally for Spider 2 refer to the Quickstart of the github page (the first point is enough)
medical_v3agentbattler-bench
AgentBattler Mini Ledger V5
Immutable evidence for 15/15 accepted Mini Ledger V5 runs across 3 harness × model conditions.
What is here
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/site/terminal-campaign.json: compact website and analysis input.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/campaign.json: source-revision-preserving campaign index with host paths removed.
snapshots/mini-ledger-v5-r5-droid-2026-08-04t13-03-57-476z/runs/:… See the full description on the dataset page: https://huggingface.co/datasets/techfren/agentbattler-bench.technical-writing-sft-100k
Technical Writing SFT (100K)
100,000 ShareGPT conversations demonstrating high-quality technical writing across 20 document types. Each example produces a complete, professional technical document — from API reference to architecture decision records to runbooks — written in the style that experienced technical writers and senior engineers actually use.
Motivation
Technical writing is one of the most underserved capabilities in LLMs. Common model failures:
Wrong… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/technical-writing-sft-100k.technicalthis is a very simple dataset i created as a test, it's not really useful for much now, but it did improve benchmarks for some older embedding models.
i was just attempting to do a quick and dirty expansion of a model's vocabulary.
it's 100% synthetic data based on lots of occupations and the tools and terms they might use in their profession.
in addition to that i added some sci-fi and fantasy terms just for laughs =)
Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques
BoxOffice Verified Seeds
This dataset contains the released BoxOffice seed datasets used in the
benchmark pipeline described in the accompanying paper. The release includes
ten verified seeds:
7
11
13
17
19
23
29
31
47
73
For each seed, we provide:
a full JSONL file containing warmup rows plus evaluation rows
an eval JSONL file containing only the evaluation rows
a manifest JSON file
a validation JSON file with directional warmup counts
Layout
viewer/
normalized… See the full description on the dataset page: https://huggingface.co/datasets/Boxoffice1280/Neurips2026_evaluating_accuracy_KV-cache_reuse_techniques.IndustryInstruction_Technology-Research
IndustryInstruction: Technology & Research
This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.forge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.Typst-Train
Typst-Train
[🤖Models] |
[🛠️Code] |
[📊Data] |
Dataset used to train Typst-Coder, includes:
18.6K Typst texts
2.5K Markdown texts containing Typst-related content
deepseek_r1_code_1kchinese-english-technical-patent-glossary
Dataset Card for 中華民國專利技術名詞中英對照詞庫
中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。
Dataset Details
Dataset Description
本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。
資料涵蓋 IPC 八大類別:
A — 人類生活需要(Human Necessities)
B — 作業、運輸(Performing Operations; Transporting)
C — 化學、冶金(Chemistry; Metallurgy)
D… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/chinese-english-technical-patent-glossary.murya-hausa-en-lexicon-robinson1914
Robinson Hausa–English Lexicon (1914)
20,628 English→Hausa word/phrase pairs, parsed from Charles Henry
Robinson's Dictionary of the Hausa Language, Volume II (English–Hausa),
3rd edition, Cambridge University Press, 1914.
This is the English→Hausa training lexicon behind
Murya — a Hausa-first voice assistant — used for
translation fine-tuning, lexical retrieval, and Hausa embedding warm-start.
Released as its own dataset so it's citable and reusable independent of the
Murya… See the full description on the dataset page: https://huggingface.co/datasets/adab-tech/murya-hausa-en-lexicon-robinson1914.kurdish-grammar-eval
Kurdish Grammar Minimal Pairs (BLiMP-style) — Kurmancî · Soranî
A grammar-competence benchmark for Kurdish, built on the BLiMP
idea: for each item, a correct sentence is paired with a corrupted version where one specific
grammar rule has been deliberately broken. Score a language model by checking whether it assigns higher
likelihood to the correct sentence than the corrupted one — accuracy well above 50% means the model
learned the rule, not just surface fluency.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/kurdish-tech/kurdish-grammar-eval.tech-qaru-tech-jobs
Russian Tech Jobs
Dataset Description
Dataset contains job vacancy posts from one of Telegram channels focusing on IT and tech recruitment.
Data Fields
id: Unique identifier of the post.
date: ISO timestamp of when the job was posted.
text: The raw text of the post with markdown.
views: The view count of the post at the time of scraping.
tags: Special tags thats will be taken from text.
How to use
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/exeyarikus/ru-tech-jobs.RPRevamped-Small
RPRevamped-Small-v1.0
Dataset Description
RPRevamped is a synthetic dataset generated by various numbers of models. It is very diverse and is recommended if you are fine-tuning a roleplay model. This is the Small version with Medium and Tiny version currently in work.
Github: RPRevamped GitHub
Here are the models used in creation of this dataset:
DeepSeek-V3-0324
Gemini-2.0-Flash-Thinking-Exp-01-21
DeepSeek-R1
Gemma-3-27B-it
Gemma-3-12B-it
Qwen2.5-VL-72B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/TechPowerB/RPRevamped-Small.bias-correction-palestine-protocol
Dataset Card for LLM Bias Correction (Palestine/Israel Context)
This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel.
Dataset Structure
The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.fake_tech_companies_market_reports
