datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tech-news-dailyCADBench-Hard
CADBench Hard Tasks
43 out of the 105 tasks. Each folder contains the complete task prompt and its authoritative Fusion reference. For all of the tasks, verifiers and sandbox environment, please reach out
Dataset categories
Domains: computer-aided design, mechanical engineering, and robotics
Modalities: natural-language task instructions and native 3D CAD artifacts
Use cases: GUI-agent evaluation, computer-use evaluation, reinforcement learning, and deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/CADBench-Hard.skill-extraction-tech
Skill Extraction with ESCO skills - TECH subset
Dataset Summary
This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-tech.us-tech-company-layoffs-warn-act-notices-daily
US tech company layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-23. 1,741 layoff notices filed by
software, IT, internet, semiconductor and electronics companies with US state
labor departments — 192,654 workers, 368
employers, 39 states, 1989–2026.
Free, CC BY 4.0, no login, no delay.
Every other tech-layoff tracker you know is assembled from news stories and
crowd submissions. This one is assembled from the legal filings employers
must make… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-tech-company-layoffs-warn-act-notices-daily.anonymous-working-histories
Structured Anonymous Career Paths extracted from Resumes
Dataset Summary
This dataset contains 2164 anonymous career paths across 24 differend industries.
Each work experience is tagger with their corresponding ESCO occupation (ESCO v1.1.1).
Languages
We use the English version of ESCO.
All resume data is in English as well.
Dataset Structure
Each working history contains up to 17 experiences.
They appear in order, and each experience has a title… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/anonymous-working-histories.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.skill-extraction-house
Skill Extraction with ESCO skills - HOUSE subset
Dataset Summary
This dataset contains an extension of the HOUSE subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-house.Synthetic-ESCO-skill-sentences
Synthetic job ads for all ESCO skills
Dataset Summary
This dataset contains 10 synthetically generated job ad sentences for almost all (99.5%) skills in ESCO v1.1.0.
Languages
We use the English version of ESCO, and all generated sentences are in English.
Dataset Structure
The dataset consists of 138,260 (sentence, skill) pairs.
Citation Information
If you use this dataset, please include the following reference:… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Synthetic-ESCO-skill-sentences.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.SkillMatch-1KWhen using this dataset, please cite:
@misc{decorte2024skillmatch,
title={SkillMatch: Evaluating Self-supervised Learning of Skill Relatedness},
author={Jens-Joris Decorte and Jeroen Van Hautte and Thomas Demeester and Chris Develder},
year={2024},
eprint={2410.05006},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.05006},
}
JobBERT-evaluation-dataset
JobBERT evaluation dataset 💾
This is the official repository containing the evaluation data that was used for the JobBERT paper. This dataset is a list of vacancy titles, each tagged with an ESCO (v1.0.5) occupation.
The full dataset is split into two files in a stratified way by class distribution. This data was automatically collected from a large governmental job board.
Access the JobBERT paper here: https://arxiv.org/abs/2109.09605
BibTeX Citation
If you use this… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/JobBERT-evaluation-dataset.techmb
Dataset Card for TechMB
Dataset Details
The Technical drawing for Manufacturability Benchmark (TechMB) gives a domain specific benchmark for the task of manufacturability evaluations based on technical drawings.
This task is described as a Visual Question Answering (VQA) task targeted at Vision Language Models (VLM) consisting of 947 question-answer pairs on 180 distinct techical drawings.
The objects, the technical drawings are developed from, represent a selection of… See the full description on the dataset page: https://huggingface.co/datasets/WSKL/techmb.span_propaganda_techniquesskill-extraction-techwolf
Skill Extraction with ESCO skills - TechWolf subset
Dataset Summary
The TECHWOLF subset, although smaller, represents a more generic distribution of job descriptions and skill spans. ESCO skills are directly annotated on the full sentence level, thus omitting the intermediate span identification step. ESCO v1.1.0 is used.
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-techwolf.bls-us-tech-employment-monthly
BLS US Tech Employment Monthly
Monthly US payroll employment for six BLS Current Employment Statistics industry series often used as a proxy for "tech employment," plus a simple monthly aggregate across those six industries.
The dataset includes:
monthly_tech_employment_components: long-form monthly data for each industry series
monthly_tech_employment_components_enriched: the same monthly component data plus a mapped top occupation for each industry
monthly_tech_employment_total:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/bls-us-tech-employment-monthly.abidhussai512_us-tech-and-ev-stock-market-dataset-2018-2026
📈 US Tech & EV Stock Market Dataset (2018 – 2026)
US Tech & EV Stock Market Dataset of AAPL, MSFT & TSLA using Python API
Dataset Info
Source: Kaggle
Original Size: 0.20 MB
Kaggle Downloads: 381
Files: 1
Files
stock_market_data.csv
Mirrored from Kaggle
cyber_MITRE_attack_tactics-and-techniquesThe dataset is question answering for MITRE tactics and techniques for version 15. Data sources are:
Tactics
Techniques
stock-technical-indicators
Stock Technical Indicators Dataset
Historical technical indicators dataset used to train directional stock movement classifiers.
Features
RSI: Relative Strength Index
SMA_20 / EMA_50: Simple and Exponential Moving Averages
MACD: Moving Average Convergence Divergence
Target: Directional label (1 = Bullish, 0 = Bearish)
press-release-benchmarks
TechBullion Press Release Builder 📰🚀
TechBullion Press Release Builder helps businesses create professional press releases, technology announcements, startup news, fintech updates, AI stories, and blockchain content ready for publication. Built by GetOnTechBullion.com.
Features
Press Release Quality Score — evaluates structure, clarity, and journalistic standards
Publication Readiness Score — checks formatting and editorial compliance
SEO Optimization Score —… See the full description on the dataset page: https://huggingface.co/datasets/get-on-techbullion/press-release-benchmarks.Tech-Stocks-NewsThis dataset contains news articles and headlines about AAPL, AMZN, TSLA, GOOG, MSFT, META, BABA, and NVDA from January 1, 2015 to January 30, 2024. The data was pulled from Alpaca Markets' API.
medium-sample-technologySample with the keyword "Technology" taken from https://huggingface.co/datasets/fabiochiu/medium-articles
TechHazardQA
Releasing our new paper How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries
👉 Read our paper at https://arxiv.org/abs/2402.15302
🎯 If you are using this dataset, please cite our papers
@article{DBLP:journals/corr/abs-2402-15302,
author = {Somnath Banerjee and
Sayan Layek and
Rima Hazra and
Animesh Mukherjee},
title = {How (un)ethical… See the full description on the dataset page: https://huggingface.co/datasets/SoftMINER-Group/TechHazardQA.tech-debt-ai-coding
Debt Behind the AI Boom — Replication Data
Data for the paper:
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo
📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding
We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five
AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static
analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.MysteryZebra
Mystery Zebra Dataset
This is the Mystery Zebra dataset created as part of the paper "Lexical Recall or Logical Reasoning: Probing the Limits of Reasoning Abilities in Large Language Models". We make the dataset available in a .csv format for your convenience. The code used to generate the puzzles in this dataset can be found in: https://github.com/arg-tech/MysteryZebra The structure of the dataset is straightforward and easy to parse. In the following, we detail the content of all… See the full description on the dataset page: https://huggingface.co/datasets/arg-tech/MysteryZebra.Aperdata-SimReady-Kitchen-01
Aperdata-SimReady-Kitchen-01
A curated collection of SimReady 3D assets for World Models and Physical AI research, especially for synthetic data generation, scene composition, embodied AI simulation, and vision-language/robot-learning experiments. Free for non-commercial use.
Dataset Version: 1.0.0
Assets list
Scene:
Kitchen_Indoor
props:
biscuit
coffee_bean
coffee_machine
container_storage
cooking_clay_pot
cooking_pot
cup_coffee
cup_general… See the full description on the dataset page: https://huggingface.co/datasets/Aperdata-tech/Aperdata-SimReady-Kitchen-01.KBYDatasetDescription
Techsalerator's KYB (Know Your Business) Data – Global Business Verification Dataset
Techsalerator's KYB (Know Your Business) Data is available for purchase by financial institutions, fintech companies, compliance and risk teams, banks, lenders, insurers, and B2B platforms worldwide. This dataset provides access to comprehensive global business data covering 400+ data points across 430M+ companies worldwide, enabling organizations to verify business entities, understand corporate… See the full description on the dataset page: https://huggingface.co/datasets/TechsaleratorLLC/KBYDataset.techchallenge-animal-condition-datasetindian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Uzaib52/indian-tech-career-intelligence-2026.peru-web-tech-stacksThis dataset contains human-verified tech stack labels for Peruvian websites based on wappalyzer and manual inspection.
frontend_layer: the library or highest level framework the client uses to render the UI.
frontend_framework: the coding framework used to structure the frontend, not necessarily SSR.
backend_framework: pure backend server or a dual stack acting only as backend.
fullstack_framework: native monolithic technologies or dual stacks such as Next.js only when they act as a… See the full description on the dataset page: https://huggingface.co/datasets/4verburga/peru-web-tech-stacks.TechnicalSupportCalls
