datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
euler-source-parquets-realeuler-source-parquetsIndustryCorpus_technology[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.forge-3b-pretrain-data
FORGE-3B Pretraining Data
Tokenized and packed pretraining data for the FORGE-3B language model.
Stats
Total tokens: 51.4070B
Domains: 10/10
Sequence length: 2048 tokens
Format: .npy shards of shape (N, 2048) with dtype uint32
Tokenizer: CRAYON (xerv-crayon, standard profile)
Domain Breakdown
Domain
Weight
Tokens (B)
Status
fineweb_edu
30%
15.0008
✓
thestack
16%
8.0011
✓
wikipedia
8%
4.2791
✓
openwebmath
8%
3.9654
✓
books
7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.world_model_tokenized_data
1X World Model Compression Challenge Dataset
This repository hosts the dataset for the 1X World Model Compression Challenge.
huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data
Updates Since v1.1
Train/Val v2.0 (~100 hours), replacing v1.1
Test v2.0 dataset for the Compression Challenge
Faces blurred for privacy
New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data
Example scripts now split into:
cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.Treble10-Speech
Treble10-Speech (16 kHz)
The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.Treble10-RIR
Treble10-RIR (32 kHz)
The Treble10-RIR dataset is a dataset for automatic speech recognition (ASR), containing high fidelity room-acoustic simulations from 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Illustrative plots of the rooms and device included in this dataset may be found in the… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-RIR.world_model_raw_dataRaw Dataset for the 1X World Model Sammpling Challenge.
Download with:
huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data
Train/Val v2.0
The training dataset is shareded into 100 independent shards. The definitions are as follows:
video_{shard}.mp4: Raw video with a resolution of 512x512.
segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.CADBench-Hard
CADBench Hard Tasks
43 out of the 105 tasks. Each folder contains the complete task prompt and its authoritative Fusion reference. For all of the tasks, verifiers and sandbox environment, please reach out
Dataset categories
Domains: computer-aided design, mechanical engineering, and robotics
Modalities: natural-language task instructions and native 3D CAD artifacts
Use cases: GUI-agent evaluation, computer-use evaluation, reinforcement learning, and deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/CADBench-Hard.claude-merged-traceseuler-structural-equationsriddle_senseriddle_sense dataset formatted into an alpaca format dataset for instruction tuning LLMs for reasoning capabilities.
Articraft-10KThis repository contains the 10k articulated 3D objects (in URDF format) from Articraft-10K.
Articraft-10K is a large-scale articulated 3D dataset generated by the Articraft agent.
IndustryCorpus2_technology_scientific_research
IndustryCorpus2: Technology & Research
This repository contains the IndustryCorpus2: Technology & Research domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_technology_scientific_research.information_technology_instruct_mcq_2481img_pointV2
img_pointV2 is available 🎉🎉🎉🥳🥳😀😀
This dataset is a collection of 3D point clouds generated from the jagennath-hari/nyuv2dataset.
img_pointV2 is the second version of the RAY-AUTRA-TECHNOLOGY/img_pointV dataset. It is a spatialized version of the NYU Depth V2 dataset, transforming classic indoor images into high-fidelity 3D point clouds (.ply files).
The main objective is to provide clean, ready-to-use 3D scenes for training 3D vision models, eliminating the need for users to… See the full description on the dataset page: https://huggingface.co/datasets/RAY-AUTRA-TECHNOLOGY/img_pointV2.librispeech_asr_sliced
Librispeech Slices
Description
Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project.
It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz.
A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment.
To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.ryan-test-white-tableThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 893,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits":{
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/ryan-test-white-table.betty-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 4,
"total_frames": 3807,
"total_tasks": 1,
"total_videos": 16,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits":{
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/betty-test.asia-science-technology-world-bank-science-and-technology-indica
Maldives - Science and Technology
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
Technological innovation, often fueled by governments, drives industrial growth and helps raise living standards. Data here aims to shed light on countries technology base: research and development, scientific and technical journal articles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-science-technology-world-bank-science-and-technology-indica.chatdoctor-embedded
Chat Doctor with Embeddings
This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped:
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
414k
Token Count
1.7b
Origin
https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view
Source of raw data
?
Processing details
paper
Embedding Model
BAAI/bge-small-en-v1.5
Data Diversity
index
Example Output
GPT-4 Rationale
GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.CircumSpectharmonia-infinity-corpusquran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.hkcancorThe Hong Kong Cantonese Corpus (HKCanCor) comprise transcribed conversations
recorded between March 1997 and August 1998. It contains recordings of
spontaneous speech (51 texts) and radio programmes (42 texts),
which involve 2 to 4 speakers, with 1 text of monologue.
In total, the corpus contains around 230,000 Chinese words.
The text is word-segmented, annotated with part-of-speech (POS) tags and
romanised Cantonese pronunciation.
Romanisation scheme - Linguistic Society of Hong Kong (LSHK)
POS scheme - Peita-Fujitsu-Renmin Ribao (PRF) corpus (Duan et al., 2000),
with extended tags for Cantonese-specific phenomena added by
Luke and Wang (see original paper for details).industrial-technical-archive
🚀 Latest Updates (July, 2026)
Version: v07.2026 (Verified)
Status: Integrated with 1,000,000+ records.
New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv.
QTE Technologies: Industrial & Scientific Knowledge Base
Wikidata Entity: Q138411149
IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq
Official Neural Hub: qtetech.github.io
This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.harmonia-triples-rust-code-traversal
harmonia-triples-rust
Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC.
Provenance
Each parquet shard carries the full provenance chain per ADR-0011:
s, p, o, src columns (when this is a triples-stage dataset)
src = "<dataset>:<version>:<file>" for triples
Causal registry events recorded at causal_registry/master.jsonl chain
Architecture
Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.touch-cokeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 10,
"total_frames": 2133,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits":{
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/touch-coke.
