CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes3.7k downloads4mo agoHugging Face02arijitghosh /T2I-ImageNet-Normalimage1M<n<10M3 likes1.3k downloads1y agoHugging Face03izumi-lab /mc4-ja-filter-ja-normal Dataset Card for "mc4-ja-filter-ja-normal" More Information needed text10M<n<100M5 likes968 downloads3y agoHugging Face04whoisandy /router-chat-normalized-1m Router Chat Normalized 1M Dataset Description Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection. Dataset Structure The dataset contains 2 split(s): train, test. Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score. Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.texttext-generation1M<n<10M0 likes838 downloads4mo agoHugging Face05izumi-lab /oscar2301-ja-filter-ja-normal Dataset Card for "oscar2301-ja-filter-ja-normal" More Information needed text10M<n<100M6 likes822 downloads3y agoHugging Face06Yanbin99 /Depth-Normal-Videos-42K Depth and Normal Videos Dataset 42,498 videos with depth and surface normals. Usage from huggingface_hub import hf_hub_download video = hf_hub_download( repo_id="Yanbin99/Depth-Normal-Videos-42K", filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4", repo_type="dataset" ) tabulardepth-estimation10K<n<100K1 likes718 downloads9mo agoHugging Face07Scicom-intl /Normalized-Multilingual-TTS Normalized Multilingual TTS Original dataset from malaysia-ai/Multilingual-TTS, we applied postfilter and postprocessing using Qwen/Qwen2.5-72B-Instruct. Acknowledgement Special thanks to https://www.scitix.ai/ for H100 Node! text10M<n<100M0 likes702 downloads6mo agoHugging Face08gebinhui /coco2017_caption_normalimage100K<n<1M0 likes595 downloads1y agoHugging Face09VidGen /Depth-Normal-Images-617Ktabular100K<n<1M0 likes534 downloads9mo agoHugging Face10N03N9 /cv24-tr-128-normalizedtext100K<n<1M0 likes528 downloads9mo agoHugging Face11N03N9 /cv24-cy-128-normalizedtext10K<n<100K0 likes474 downloads9mo agoHugging Face12winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-topk128-normalizedtabular10K<n<100K1 likes437 downloads1y agoHugging Face13N03N9 /cv24-ur-128-normalizedtext10K<n<100K0 likes399 downloads9mo agoHugging Face14N03N9 /cv24-pt-128-normalizedtext100K<n<1M0 likes398 downloads9mo agoHugging Face15N03N9 /cv24-uk-128-normalizedtext10K<n<100K0 likes398 downloads9mo agoHugging Face16N03N9 /cv24-sw-128-normalizedtext100K<n<1M0 likes398 downloads9mo agoHugging Face17N03N9 /cv24-de-128-normalizedtext100K<n<1M0 likes392 downloads9mo agoHugging Face18N03N9 /cv24-sk-128-normalizedtext10K<n<100K0 likes380 downloads9mo agoHugging Face19mateuszgrzyb /lichess-stockfish-normalized Lichess Chess Positions: ML-Ready Deduplicated Evaluations Dataset Description A curated dataset of 316,072,343 unique chess positions with Stockfish evaluations, optimized for training neural networks. This is a deduplicated, ML-ready version of the Lichess evaluation database. Why This Dataset? While Lichess provides deduplicated evaluations in JSONL.zst format, and HuggingFace hosts the full (non-deduplicated) version, this dataset offers: Unique advantages:… See the full description on the dataset page: https://huggingface.co/datasets/mateuszgrzyb/lichess-stockfish-normalized.tabulartabular-regression100M<n<1B4 likes366 downloads10mo agoHugging Face20ChenWu98 /stack-v2-python-normal-onlytext1M<n<10M0 likes360 downloads1y agoHugging Face21winglian /OpenThoughts-114k-math-correct-qwen3-14b-math-prepared-normalizedtabular10K<n<100K0 likes323 downloads1y agoHugging Face22Pointcept /modelnet40_normal_resampled-compressedtext10K<n<100K2 likes314 downloads2y agoHugging Face23N03N9 /cv24-ca-128-normalizedtext1M<n<10M0 likes307 downloads9mo agoHugging Face24ducido /merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_taskThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "panda", "total_episodes": 800, "total_frames": 133851, "total_tasks": 40, "total_videos": 0, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:800" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ducido/merged_libero_s40_mdnl_original_s40_synthetic_20ep_2_hard_task_5ep_all_normal_task.imagerobotics100K<n<1M0 likes298 downloads6mo agoHugging Face25csoai /gspc-normalized GSPC normalised — every bank in one schema The one schema to read first. 518 rows that flatten several GSPC banks into a single shape: source (the bank repository the row came from), axis, category, anchor, prompt, expected, expected_is_list, and raw_keys recording the original row's keys so nothing is silently dropped. If you want to reuse the banks without learning each one's native layout, start here. The live board is the authority GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-normalized.textothern<1K0 likes240 downloads8d agoHugging Face26TechWolf /Skill-normalisation-ESCO-graded skill-normalisation-esco-graded Graded-relevance annotations for surface skill terms (ESCO alt-labels) from ESCO v1.1.0 skill-normalisation pairs against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 50 _id (query id), text (ESCO alt-label / surface term to… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-normalisation-ESCO-graded.text1M<n<10M0 likes236 downloads1mo agoHugging Face27Twelve2five /igbo_tts_normalizedaudio100K<n<1M2 likes227 downloads1y agoHugging Face28v1v1d /v1v1d_docmatix_1k_normal_v6image10K<n<100K0 likes226 downloads8mo agoHugging Face29v1v1d /v1v1d_docmatix_1k_normal_v5image10K<n<100K0 likes213 downloads8mo agoHugging Face30Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes203 downloads16d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.