datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anonymous-working-histories
Structured Anonymous Career Paths extracted from Resumes
Dataset Summary
This dataset contains 2164 anonymous career paths across 24 differend industries.
Each work experience is tagger with their corresponding ESCO occupation (ESCO v1.1.1).
Languages
We use the English version of ESCO.
All resume data is in English as well.
Dataset Structure
Each working history contains up to 17 experiences.
They appear in order, and each experience has a title… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/anonymous-working-histories.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.techmb
Dataset Card for TechMB
Dataset Details
The Technical drawing for Manufacturability Benchmark (TechMB) gives a domain specific benchmark for the task of manufacturability evaluations based on technical drawings.
This task is described as a Visual Question Answering (VQA) task targeted at Vision Language Models (VLM) consisting of 947 question-answer pairs on 180 distinct techical drawings.
The objects, the technical drawings are developed from, represent a selection of… See the full description on the dataset page: https://huggingface.co/datasets/WSKL/techmb.bls-us-tech-employment-monthly
BLS US Tech Employment Monthly
Monthly US payroll employment for six BLS Current Employment Statistics industry series often used as a proxy for "tech employment," plus a simple monthly aggregate across those six industries.
The dataset includes:
monthly_tech_employment_components: long-form monthly data for each industry series
monthly_tech_employment_components_enriched: the same monthly component data plus a mapped top occupation for each industry
monthly_tech_employment_total:… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/bls-us-tech-employment-monthly.stock-technical-indicators
Stock Technical Indicators Dataset
Historical technical indicators dataset used to train directional stock movement classifiers.
Features
RSI: Relative Strength Index
SMA_20 / EMA_50: Simple and Exponential Moving Averages
MACD: Moving Average Convergence Divergence
Target: Directional label (1 = Bullish, 0 = Bearish)
press-release-benchmarks
TechBullion Press Release Builder 📰🚀
TechBullion Press Release Builder helps businesses create professional press releases, technology announcements, startup news, fintech updates, AI stories, and blockchain content ready for publication. Built by GetOnTechBullion.com.
Features
Press Release Quality Score — evaluates structure, clarity, and journalistic standards
Publication Readiness Score — checks formatting and editorial compliance
SEO Optimization Score —… See the full description on the dataset page: https://huggingface.co/datasets/get-on-techbullion/press-release-benchmarks.tech-debt-ai-coding
Debt Behind the AI Boom — Replication Data
Data for the paper:
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo
📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding
We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five
AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static
analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Uzaib52/indian-tech-career-intelligence-2026.KBYDatasetDescription
Techsalerator's KYB (Know Your Business) Data – Global Business Verification Dataset
Techsalerator's KYB (Know Your Business) Data is available for purchase by financial institutions, fintech companies, compliance and risk teams, banks, lenders, insurers, and B2B platforms worldwide. This dataset provides access to comprehensive global business data covering 400+ data points across 430M+ companies worldwide, enabling organizations to verify business entities, understand corporate… See the full description on the dataset page: https://huggingface.co/datasets/TechsaleratorLLC/KBYDataset.tech-speech-congress
Tech-Speech in the Congressional Record, 1995–2025
Descriptive aggregates from a speech-level analysis of technology-related
discourse in the U.S. Congressional Record, 1995–2025. Built from
govinfo.gov daily Congressional Record packages, parsed to the individual
speech turn with speaker metadata (party, state, chamber, Bioguide ID),
filtered for procedural speech, and scored with a TF-IDF technology-intensity
index over a 237-term technology vocabulary built from federal and… See the full description on the dataset page: https://huggingface.co/datasets/tapanyemre/tech-speech-congress.bitcoin-ethereum-orderflow-cvd-alpha
Bitcoin & Ethereum 1-Minute Order Flow & Cumulative Volume Delta (CVD) Alpha
Institutional Market Microstructure Dataset Sample (Clean CSV / Parquet Ready)
📌 Dataset Overview
In cryptocurrency and traditional electronic markets, price action is driven by aggressive market orders (taker flow) that cross the bid-ask spread. This preview dataset provides 1,000 rows of continuous 1-minute order flow for Bitcoin (BTC/USDT) and Ethereum (ETH/USDT)… See the full description on the dataset page: https://huggingface.co/datasets/TechPlayground/bitcoin-ethereum-orderflow-cvd-alpha.Scientific-and-technical-journal-articles-Africa
Scientific and technical journal articles Africa | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Scientific-and-technical-journal-articles-Africa.difficult-technology-8dac0d
difficult-technology-8dac0d
Synthetic products test data: 56 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Craft89/difficult-technology-8dac0d.africa-synth-education-technical-vocational-training-comoros
Africa Synth Education Technical Vocational Training Comoros | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-technical-vocational-training-comoros.House-Prices-Advanced-Regression-Techniques-result-0.12387abdullahkhan70_github-tech-stack-languages-and-frameworks
GitHub Tech Stack Languages & Frameworks
Comprehensive Repository Data: JavaScript, Python, Go, Rust & More
Dataset Info
Source: Kaggle
Original Size: 2.17 MB
Kaggle Downloads: 62
Files: 17
Files
Mirrored from Kaggle
techmb
Dataset Card for TechMB
Dataset Details
The Technical drawing for Manufacturability Benchmark (TechMB) gives a domain specific benchmark for the task of manufacturability evaluations based on technical drawings.
This task is described as a Visual Question Answering (VQA) task targeted at Vision Language Models (VLM) consisting of 947 question-answer pairs on 180 distinct techical drawings.
The objects, the technical drawings are developed from, represent a… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/techmb.Technicians-in-Research-and-Development-in-Africa-per-million-people
Technicians in Research and Development in Africa per million people | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Technicians-in-Research-and-Development-in-Africa-per-million-people.High-technology-exports-current-USD-africa
High technology exports current USD africa | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/High-technology-exports-current-USD-africa.BertHigh-technology-exports-current-US-Dollars
High technology exports current US Dollars | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: technology_digital - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/High-technology-exports-current-US-Dollars.indian-tech-career-intelligence-2026
India Tech Career Intelligence [1M]
About Dataset
India Tech Career Intelligence [1M] is a comprehensive, production-grade dataset containing 1,000,000 (1 Million) standardized records representing the Indian technology job and internship ecosystem.
The dataset has been designed for Data Scientists, Machine Learning Engineers, Analysts, Researchers, Students, and Developers interested in understanding hiring trends, salary distributions, skill demand, and… See the full description on the dataset page: https://huggingface.co/datasets/Jidnesh298/indian-tech-career-intelligence-2026.techafrica-synth-education-technology-adoption-africa-all
Africa Synth Education Technology Adoption Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-technology-adoption-africa-all.tech-ready-restaurants-in-the-boston-cambridge-newton-metro-area-ma-nh-us-160567
Tech-Ready Restaurants in the Boston-Cambridge-Newton Metro Area, MA-NH, US
Free sample dataset from BeamStation
The dataset "Tech-Ready Restaurants in the Boston-Cambridge-Newton Metro Area, MA-NH, US" lists 49 establishments that meet specific criteria for technology adoption. These restaurants are well‑established, indicated by a Beam Score above 70, and show recent positive performance with sentiment scores over 10 in the last 30 days. Despite their strong foot traffic and… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/tech-ready-restaurants-in-the-boston-cambridge-newton-metro-area-ma-nh-us-160567.tech_reports_miningaxya-tech-websearch
Dhivehi Combined Dataset
Overview
This dataset combines three 30K Dhivehi language corpora from the Leipzig Corpora Collection into a single unified CSV file containing 90,000 sentences. The dataset provides a comprehensive resource for Dhivehi language processing, combining data from Wikipedia, news sources, and web crawls.
Dataset Composition
The dataset consists of three distinct sources:
Wikipedia (2021): 30,000 sentences from Dhivehi Wikipedia
News… See the full description on the dataset page: https://huggingface.co/datasets/axmeeabdhullo/axya-tech-websearch.tech-ready-restaurants-in-the-chicago-naperville-elgin-metro-area-il-in-us-171859
Tech-Ready Restaurants in the Chicago-Naperville-Elgin Metro Area, IL-IN, US
Free sample dataset from BeamStation
The "Tech-Ready Restaurants in the Chicago-Naperville-Elgin Metro Area, IL-IN, US" dataset highlights a prime segment of dining establishments that are financially solid but technologically underserved. Covering the Chicago-Naperville-Elgin metropolitan region, it contains 105 records of restaurants that score above 70 on the Beam Score, indicating strong establishment… See the full description on the dataset page: https://huggingface.co/datasets/beamstation/tech-ready-restaurants-in-the-chicago-naperville-elgin-metro-area-il-in-us-171859.Medium-and-high-tech-exports-percentage-of-manufactured-exports
Medium and high tech exports percentage of manufactured exports | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Medium-and-high-tech-exports-percentage-of-manufactured-exports.
