CoolFace
20 results

techno

Hailstone-Technologies /euler-source-parquets-realtext1M<n<10M0 likes4k downloads3mo agoHugging FaceHailstone-Technologies /euler-source-parquetstext1M<n<10M0 likes4k downloads3mo agoHugging FaceBAAI /IndustryCorpus_technology[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_technology.texttext-generation10M<n<100M4 likes3.7k downloads1mo agoHugging FacePhase-Technologies /forge-3b-pretrain-data FORGE-3B Pretraining Data Tokenized and packed pretraining data for the FORGE-3B language model. Stats Total tokens: 51.4070B Domains: 10/10 Sequence length: 2048 tokens Format: .npy shards of shape (N, 2048) with dtype uint32 Tokenizer: CRAYON (xerv-crayon, standard profile) Domain Breakdown Domain Weight Tokens (B) Status fineweb_edu 30% 15.0008 ✓ thestack 16% 8.0011 ✓ wikipedia 8% 4.2791 ✓ openwebmath 8% 3.9654 ✓ books 7%… See the full description on the dataset page: https://huggingface.co/datasets/Phase-Technologies/forge-3b-pretrain-data.text-generation10B<n<100B0 likes3.6k downloads3mo agoHugging FaceQUD-Technologies /quranic-universal-ayahs Qur'anic Universal Ayahs Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset. This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.audioautomatic-speech-recognition100K<n<1M6 likes2.8k downloads2d agoHugging Face1x-technologies /world_model_tokenized_data 1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.10M<n<100M34 likes2.7k downloads1y agoHugging Face