datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube_technical_v0
youtube_technical_v0
Canonical Phonon ASR dataset. Audio is embedded in Parquet as a Hugging
Face-compatible audio struct with bytes and path.
Hidden eval rows must not be used for training or synthetic prompt generation.
Hub
Dataset id: Infatoshi/youtube_technical_v0 (public)
Companion labels/manifests/attribution: Infatoshi/phonon-youtube-technical
Rows: see summary.json (~118k)
Audio is embedded in Parquet (audio struct: bytes, path)
Split… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/youtube_technical_v0.baguwen-technical-qa
八股文技术面试题库 · Baguwen Technical Interview QA
一套面向中文技术八股/面试的问答数据集,收集并整理自三个公开 GitHub 仓库,统一为 (问题, 答案) 结构化格式,适合检索增强(RAG)、SFT 微调、面试题库等场景。
数据来源(已标注)
来源
仓库
类型
处理方式
源头更新时间
状态
bestJavaer
crisxuan/bestJavaer
手写 Markdown 笔记(168 篇)
直接解析原始 Markdown(## 问题 + 答案),按篇拆分,并剔除作者主观/营销干扰文本
2026-07-28
✅ 已收录(631 条)
learning_mind_map
0voice/learning_mind_map
扫描版思维导图 PDF(74 个,已成功转换 49 个)
PyMuPDF 渲染每页为图片 → qwen3.8-27b 视觉 OCR → 层级 Markdown → LLM 改写为问答对
2024-05-20
🟡 部分收录(198 条;剩余… See the full description on the dataset page: https://huggingface.co/datasets/Weidows/baguwen-technical-qa.technical-manuals
Description
Topic: Technical Manuals
Domains: Engineering, Information Technology, Product Documentation
Number of Entries: 1,000
Dataset Type: Raw Dataset
Model Used: bedrock/us.meta.llama4-maverick-17b-instruct-v1:0
Language: English
depthapi_technical_corpus
DepthAPI Technical Corpus
Overview
The DepthAPI Technical Corpus is a curated, high-quality retrieval corpus designed for modern RAG (Retrieval-Augmented Generation) systems. It features clean, aggressively normalized technical documentation, code snippets, engineering post-mortems, and system design literature.
This dataset was explicitly built to serve as the local ground-truth for the DepthAPI project.
About the DepthAPI Project
DepthAPI is an… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/depthapi_technical_corpus.ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.stk-technical-indicators-1hourafrica-synth-education-vocational-technical-nigeria
Africa Synth Education Vocational Technical Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-vocational-technical-nigeria.Multimodal-test-dataset-technicalindicators
KRX Investment Warning Prediction Dataset (OHLCV + Technical Indicators + Korean News)
Dataset Summary
This dataset is a test dataset for predicting Investment Warning (투자주의종목) designations in the Korean stock market (KRX).
It contains raw daily OHLCV price data, 13 technical indicators, and Korean news text (title + body), designed for multimodal anomaly detection / binary classification.
Important: No normalization/scaling is applied. All values are raw.
Date range:… See the full description on the dataset page: https://huggingface.co/datasets/k-datasoft/Multimodal-test-dataset-technicalindicators.mulesoft-technical-articles
SkillPilot Weaviate RAG Dataset
This dataset is an export of the SkillPilot RAG (Retrieval-Augmented Generation) knowledge base, clustered and enhanced with Self-Organizing Maps (SOM), and stored in a Weaviate vector database. The export is provided in Parquet format for efficient analysis and machine learning workflows.
Dataset Overview
Source: SkillPilot Weaviate vector database
Export Date: July 8, 2025
Format: Parquet (with additional JSON stats)
Total Chunks: 11,412… See the full description on the dataset page: https://huggingface.co/datasets/BassemE/mulesoft-technical-articles.TTS_English_Technical_Termstechnical-support-dataset
Dataset Card for technical-support-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/hk-gaianet/technical-support-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/harishkotra/technical-support-dataset.TTS_English_Technical_datastk-technical-indicators-15minasia-owid-total-official-development-assistance-for-technical-cooperation
Total Official Development Assistance For Technical Cooperation | Asia (Our World in Data)
🌏 888 observations · 39 Asia countries · 2000–2023 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 888 observations of Total Official Development Assistance For Technical Cooperation data across 39 Asia countries, spanning 2000–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Total… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-total-official-development-assistance-for-technical-cooperation.europe-owid-total-official-development-assistance-for-technical-cooperation
Total Official Development Assistance For Technical Cooperation | Europe (Our World in Data)
🇪🇺 173 observations · 10 Europe countries · 2000–2023 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 173 observations of Total Official Development Assistance For Technical Cooperation data across 10 Europe countries, spanning 2000–2023.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-total-official-development-assistance-for-technical-cooperation.eng_latn_code_technical_content
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_code_technical_content.hawk-technical-memoryafrica-worldbank-world-bank-share-of-total-education-expenditure-for-technical-vocational-educat
World Bank: Share of total education expenditure for technical/vocational education (%) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-world-bank-share-of-total-education-expenditure-for-technical-vocational-educat.stk-technical-indicators-1minelevenlabs_multilingual_v2-technical-speech
ElevenLabs Multilingual V2 Technical Speech Dataset
This dataset contains automatically generated technical phrases in three domains, converted to speech using the ElevenLabs Multilingual V2 model with Adam voice.
Dataset Description
The dataset includes audio samples of technical phrases across three categories:
Machine Learning (ML)
Science
Technology
Each entry contains:
Audio file in MP3 format (22050Hz)
Source text
Text length
Category label
Data… See the full description on the dataset page: https://huggingface.co/datasets/WpythonW/elevenlabs_multilingual_v2-technical-speech.eng_latn_code_technical_content
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/eng_latn_code_technical_content.code_technical_content
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/code_technical_content.ray-technical-docs-qa-chunks
Dataset Card for "final_data"
More Information needed
TechnicalDatastk-technical-indicators-5minfx-technical-indicators-30minstk-technical-indicators-30minafrica-worldbank-vocational-and-technical-enrolment-of-total-secondary-enrolment-female-se-sec-e
Vocational and Technical enrolment (% of total secondary enrolment), female | Africa (World Bank — Gender Statistics) | Africa (World Bank)
Size category: 1K<n<10K - Formats: parquet - Sector: demographics_social - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-vocational-and-technical-enrolment-of-total-secondary-enrolment-female-se-sec-e.stk-technical-indicators-4hourmy_technical_support_chatbot_dataset
