CoolFace
20 results

normalized

mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes3.7k downloads4mo agoHugging FaceSaakethS /fishnet-lichess-normalized1 likes2.1k downloads8d agoHugging FaceTongheZhangTH /XDof-TshirtFolding-20hours-normalizedtabular1M<n<10M0 likes1.4k downloads8mo agoHugging Facewhoisandy /router-chat-normalized-1m Router Chat Normalized 1M Dataset Description Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection. Dataset Structure The dataset contains 2 split(s): train, test. Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score. Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.texttext-generation1M<n<10M0 likes838 downloads4mo agoHugging Facecfy2yue /perturbseq_normalized CellClip public normalized perturbation single-cell release This public repository contains the source-cleared, human, cell-level portion of CellClip Stage 1: 254 H5AD files, 12,718,270 cells, and 260,859,318,437 payload bytes. These are processed derivatives rather than the original raw download archives. Source partition Units Cells Reference cells Effect cells scPerturb (23 genepert + 229 chempert) 252 4,774,802 397,528 4,377,274 XCell / X-Atlas Orion (genepert)… See the full description on the dataset page: https://huggingface.co/datasets/cfy2yue/perturbseq_normalized.feature-extraction0 likes800 downloads2mo agoHugging FaceScicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes702 downloads6mo agoHugging Face