datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
euler-source-parquetsIndustryCorpus_technology[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.euler-source-parquets-realquranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.Treble10-Speech
Treble10-Speech (16 kHz)
The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.Treble10-RIR
Treble10-RIR (32 kHz)
The Treble10-RIR dataset is a dataset for automatic speech recognition (ASR), containing high fidelity room-acoustic simulations from 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Illustrative plots of the rooms and device included in this dataset may be found in the… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-RIR.CADBench-Hard
CADBench Hard Tasks
43 out of the 105 tasks. Each folder contains the complete task prompt and its authoritative Fusion reference. For all of the tasks, verifiers and sandbox environment, please reach out
Dataset categories
Domains: computer-aided design, mechanical engineering, and robotics
Modalities: natural-language task instructions and native 3D CAD artifacts
Use cases: GUI-agent evaluation, computer-use evaluation, reinforcement learning, and deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/CADBench-Hard.euler-structural-equationsclaude-merged-tracesriddle_senseriddle_sense dataset formatted into an alpaca format dataset for instruction tuning LLMs for reasoning capabilities.
information_technology_instruct_mcq_2481librispeech_asr_sliced
Librispeech Slices
Description
Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project.
It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz.
A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment.
To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.quran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.chatdoctor-embedded
Chat Doctor with Embeddings
This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped:
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
414k
Token Count
1.7b
Origin
https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view
Source of raw data
?
Processing details
paper
Embedding Model
BAAI/bge-small-en-v1.5
Data Diversity
index
Example Output
GPT-4 Rationale
GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.asia-science-technology-world-bank-science-and-technology-indica
Maldives - Science and Technology
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
Technological innovation, often fueled by governments, drives industrial growth and helps raise living standards. Data here aims to shed light on countries technology base: research and development, scientific and technical journal articles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-science-technology-world-bank-science-and-technology-indica.harmonia-infinity-corpusharmonia-triples-rust-code-traversal
harmonia-triples-rust
Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC.
Provenance
Each parquet shard carries the full provenance chain per ADR-0011:
s, p, o, src columns (when this is a triples-stage dataset)
src = "<dataset>:<version>:<file>" for triples
Causal registry events recorded at causal_registry/master.jsonl chain
Architecture
Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.africa-owid-access-to-clean-fuels-and-technologies-for-cooking
Access To Clean Fuels And Technologies For Cooking | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-access-to-clean-fuels-and-technologies-for-cooking.IndustryInstruction_Technology-Research
IndustryInstruction: Technology & Research
This repository contains the IndustryInstruction: Technology & Research domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Technology-Research.harmonia-triples-stackexchange-document-traversal
harmonia-triples-stackexchange-slice
Triples for source stackexchange-slice emitted by the ingest pipeline (current wave: v0.6). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC.
Provenance
Each parquet shard carries the full provenance chain per ADR-0011:
s, p, o, src columns (when this is a triples-stage dataset)
src = "<dataset>:<version>:<file>" for triples
Causal registry events recorded at causal_registry/master.jsonl chain… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-stackexchange-document-traversal.forge-3b-dpo-data
FORGE-3B DPO Preference Data
Tokenized (prompt, chosen, rejected) preference triples for DPO post-training
of FORGE-3B, built per the FORGE paper Section 6.2 / Appendix A.2.
This is data preparation output only — no model was trained to produce this.
Stats
Total pairs: 0 (paper target: ~200,000)
Domains: 0/4
Context length: 4096 tokens (paper Appendix A.2, DPO block)
Format: unpacked — one (prompt, chosen, rejected) triple per training example
Chat template:… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-dpo-data.MedCase-Structured
MedCase-Structured
Dataset for Paper MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings
Structured FHIR R4 representations of clinical reasoning cases, derived from the
MedCaseReasoning dataset (Wu et al., 2025). Each case pairs a free-text
clinical presentation with a machine-readable FHIR bundle and a held-out
ground-truth diagnosis, supporting evaluation of clinical information
extraction, terminology coding… See the full description on the dataset page: https://huggingface.co/datasets/system-technologies/MedCase-Structured.harmonia-graph-causal
harmonia-graph-causal
Causal graph triples from structured sources: bnlearn Bayesian networks, Reactome pathways, STRING protein interactions, SIGNOR signaling, Wikidata causal properties, ConceptNet causal relations, Tübingen cause-effect pairs, and 60+ Harmonia Structural Causal Models covering the full human experience.
Schema
Every row: s (subject), p (predicate), o (object), src (provenance)
src format: dataset:version:file
Part of the Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-graph-causal.synthetic-clinical-notes-embedded
Synthetic Clinical Notes
This dataset is post-processed version of starmpcc/Asclepius-Synthetic-Clinical-Notes:
Turn into Alpaca format (instruction, input, and output)
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
158k
Token Count
648m
Origin
https://figshare.com/authors/Zhengyun_Zhao/16480335
Source of raw data
PubMed Central (PMC) and MIMIC 3
Processing details
original, paper
Embedding Model… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/synthetic-clinical-notes-embedded.gold-reddit-ai-technology-corpus
Gold Reddit AI and Technology Corpus
Dataset Authors
Ricardo Flores-Moyano, Universidad San Francisco de Quito
Felipe Rosero-Polo, Universidad San Francisco de Quito
José Vega-Sánchez, Universidad San Francisco de Quito
Maria Baldeon-Calisto, Wake Forest University
Dataset Summary
This dataset provides a curated Gold corpus of 1,614 Reddit posts that were
publicly accessible at collection time. The corpus is primarily English, with
a smaller… See the full description on the dataset page: https://huggingface.co/datasets/rickphd/gold-reddit-ai-technology-corpus.medium-sample-technologySample with the keyword "Technology" taken from https://huggingface.co/datasets/fabiochiu/medium-articles
medical-prescriptionsLab-1-language-technology
