CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hailstone-Technologies /euler-source-parquets-realtext1M<n<10M0 likes4k downloads3mo agoHugging Face02Hailstone-Technologies /euler-source-parquetstext1M<n<10M0 likes4k downloads3mo agoHugging Face03BAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.7k downloads1mo agoHugging Face04Phase-Technologies /forge-3b-pretrain-data FORGE-3B Pretraining Data Tokenized and packed pretraining data for the FORGE-3B language model. Stats Total tokens: 51.4070B Domains: 10/10 Sequence length: 2048 tokens Format: .npy shards of shape (N, 2048) with dtype uint32 Tokenizer: CRAYON (xerv-crayon, standard profile) Domain Breakdown Domain Weight Tokens (B) Status fineweb_edu 30% 15.0008 ✓ thestack 16% 8.0011 ✓ wikipedia 8% 4.2791 ✓ openwebmath 8% 3.9654 ✓ books 7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.text-generation10B<n<100B0 likes3.6k downloads3mo agoHugging Face05QUD-Technologies /quranic-universal-ayahs Qur'anic Universal Ayahs Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset. This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.audioautomatic-speech-recognition100K<n<1M6 likes2.8k downloads2d agoHugging Face061x-technologies /world_model_tokenized_data 1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.10M<n<100M34 likes2.7k downloads1y agoHugging Face07TechnoBaptist /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.tabulartext-generation100M<n<1B0 likes1.2k downloads2mo agoHugging Face08treble-technologies /Treble10-Speech Treble10-Speech (16 kHz) The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms. The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s. Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.audioautomatic-speech-recognition1K<n<10K21 likes1.1k downloads11mo agoHugging Face09treble-technologies /Treble10-RIR Treble10-RIR (32 kHz) The Treble10-RIR dataset is a dataset for automatic speech recognition (ASR), containing high fidelity room-acoustic simulations from 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms. The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s. Illustrative plots of the rooms and device included in this dataset may be found in the… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-RIR.audio1K<n<10K33 likes863 downloads11mo agoHugging Face101x-technologies /world_model_raw_dataRaw Dataset for the 1X World Model Sammpling Challenge. Download with: huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data Train/Val v2.0 The training dataset is shareded into 100 independent shards. The definitions are as follows: video_{shard}.mp4: Raw video with a resolution of 512x512. segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.10M<n<100M8 likes621 downloads1y agoHugging Face11Seldon-Technologies /CADBench-Hard CADBench Hard Tasks 43 out of the 105 tasks. Each folder contains the complete task prompt and its authoritative Fusion reference. For all of the tasks, verifiers and sandbox environment, please reach out Dataset categories Domains: computer-aided design, mechanical engineering, and robotics Modalities: natural-language task instructions and native 3D CAD artifacts Use cases: GUI-agent evaluation, computer-use evaluation, reinforcement learning, and deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Seldon-Technologies/CADBench-Hard.textreinforcement-learningn<1K0 likes564 downloads1mo agoHugging Face12Phase-Technologies /claude-merged-tracestext100K<n<1M0 likes466 downloads2mo agoHugging Face13Hailstone-Technologies /euler-structural-equationstabular10M<n<100M0 likes434 downloads3mo agoHugging Face14Technoculture /riddle_senseriddle_sense dataset formatted into an alpaca format dataset for instruction tuning LLMs for reasoning capabilities. textquestion-answering1K<n<10K1 likes413 downloads3y agoHugging Face15TechnoBaptist /Articraft-10KThis repository contains the 10k articulated 3D objects (in URDF format) from Articraft-10K. Articraft-10K is a large-scale articulated 3D dataset generated by the Articraft agent. 0 likes412 downloads2mo agoHugging Face16BAAI /IndustryCorpus2_technology_scientific_research IndustryCorpus2: Technology & Research This repository contains the IndustryCorpus2: Technology & Research domain subset of BAAI/IndustryCorpus2. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryCorpus2: @misc{shi2024industrycorpus2, title = {IndustryCorpus2}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_technology_scientific_research.1 likes367 downloads1mo agoHugging Face17gussieIsASuccessfulWarlock /information_technology_instruct_mcq_2481textn<1K2 likes365 downloads2y agoHugging Face18RAY-AUTRA-TECHNOLOGY /img_pointV2 img_pointV2 is available 🎉🎉🎉🥳🥳😀😀 This dataset is a collection of 3D point clouds generated from the jagennath-hari/nyuv2dataset. img_pointV2 is the second version of the RAY-AUTRA-TECHNOLOGY/img_pointV dataset. It is a spatialized version of the NYU Depth V2 dataset, transforming classic indoor images into high-fidelity 3D point clouds (.ply files). The main objective is to provide clean, ready-to-use 3D scenes for training 3D vision models, eliminating the need for users to… See the full description on the dataset page: https://huggingface.co/datasets/RAY-AUTRA-TECHNOLOGY/img_pointV2.3d1 likes339 downloads9mo agoHugging Face19treble-technologies /librispeech_asr_sliced Librispeech Slices Description Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project. It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz. A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment. To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.tabular100K<n<1M0 likes285 downloads7mo agoHugging Face20allday-technology /ryan-test-white-tableThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_stationary", "total_episodes": 1, "total_frames": 893, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits":{ "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/ryan-test-white-table.tabularrobotics100K<n<1M0 likes265 downloads11mo agoHugging Face21allday-technology /betty-testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_stationary", "total_episodes": 4, "total_frames": 3807, "total_tasks": 1, "total_videos": 16, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits":{ "train": "0:4" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/betty-test.tabularrobotics10K<n<100K0 likes247 downloads1y agoHugging Face22electricsheepasia /asia-science-technology-world-bank-science-and-technology-indica Maldives - Science and Technology Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28 Abstract Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX. Technological innovation, often fueled by governments, drives industrial growth and helps raise living standards. Data here aims to shed light on countries technology base: research and development, scientific and technical journal articles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-science-technology-world-bank-science-and-technology-indica.tabulartabular-classificationn<1K0 likes244 downloads5mo agoHugging Face23Technoculture /chatdoctor-embedded Chat Doctor with Embeddings This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped: Add embeddings for input and output columns using BAAI/bge-small-en-v1.5 Details Sample Count 414k Token Count 1.7b Origin https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view Source of raw data ? Processing details paper Embedding Model BAAI/bge-small-en-v1.5 Data Diversity index Example Output GPT-4 Rationale GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.text100K<n<1M3 likes236 downloads3y agoHugging Face24crystal-technologies /CircumSpect0 likes229 downloads2y agoHugging Face25Hailstone-Technologies /harmonia-infinity-corpustext100M<n<1B0 likes228 downloads5mo agoHugging Face26QUD-Technologies /quran-alignment-benchmark Quran Recitation Alignment Benchmark Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here. 16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.audioautomatic-speech-recognitionn<1K0 likes224 downloads14d agoHugging Face27nanyang-technological-university-singapore /hkcancorThe Hong Kong Cantonese Corpus (HKCanCor) comprise transcribed conversations recorded between March 1997 and August 1998. It contains recordings of spontaneous speech (51 texts) and radio programmes (42 texts), which involve 2 to 4 speakers, with 1 text of monologue. In total, the corpus contains around 230,000 Chinese words. The text is word-segmented, annotated with part-of-speech (POS) tags and romanised Cantonese pronunciation. Romanisation scheme - Linguistic Society of Hong Kong (LSHK) POS scheme - Peita-Fujitsu-Renmin Ribao (PRF) corpus (Duan et al., 2000), with extended tags for Cantonese-specific phenomena added by Luke and Wang (see original paper for details).translation10K<n<100K17 likes212 downloads3y agoHugging Face28QTE-Technologies /industrial-technical-archive 🚀 Latest Updates (July, 2026) Version: v07.2026 (Verified) Status: Integrated with 1,000,000+ records. New Files: product-E-26-07-2026.csv & product-V-26-07-2026.csv. QTE Technologies: Industrial & Scientific Knowledge Base Wikidata Entity: Q138411149 IPFS CID: bafybeibogxxuhmzfrsuhcfd4qr4tmc4okhmrcwhp3266hq47ccuyjnjxoq Official Neural Hub: qtetech.github.io This is the permanent technical archive for QTE Technologies, ensuring long-term accessibility of… See the full description on the dataset page: https://huggingface.co/datasets/QTE-Technologies/industrial-technical-archive.image10K<n<100K0 likes204 downloads2mo agoHugging Face29Hailstone-Technologies /harmonia-triples-rust-code-traversal harmonia-triples-rust Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC. Provenance Each parquet shard carries the full provenance chain per ADR-0011: s, p, o, src columns (when this is a triples-stage dataset) src = "<dataset>:<version>:<file>" for triples Causal registry events recorded at causal_registry/master.jsonl chain Architecture Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.textgraph-ml1M<n<10M0 likes199 downloads5mo agoHugging Face30allday-technology /touch-cokeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "trossen_subversion": "v1.0", "robot_type": "trossen_ai_stationary", "total_episodes": 10, "total_frames": 2133, "total_tasks": 1, "total_videos": 40, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits":{ "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/touch-coke.tabularrobotics10K<n<100K0 likes193 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.