CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZhejiangLab /CPT_Data_Pool CPT Data Pool This repository hosts a large-scale CPT corpus designed for continual pre-training of domain-specific large language models. It serves as a data component of the 👉F2D-LLM framework, an end-to-end pipeline for domain-specific LLM training. For detailed data processing, scoring methods, and sampling strategies, please refer to the official 👉GitHub repo. Overall, the dataset contains approximately 271B tokens and is stored as jsonl files with a unified schema… See the full description on the dataset page: https://huggingface.co/datasets/ZhejiangLab/CPT_Data_Pool.text100M<n<1B1 likes4.8k downloads3mo agoHugging Face02jiviteshjn /fineweb-edu-zh-chengyu-cpt Fineweb-Edu Chinese — Chengyu-Tagged Continued-Pretraining Corpus A 3.74M-document Chinese corpus (~7.8B tokens) for continued pretraining on cultural knowledge in figurative language, built from the highest-quality tier of opencsg/Fineweb-Edu-Chinese-V2.1. Each document is educational Chinese text containing at least one culturally vetted chengyu, with an appended 【成语注释】 knowledge block listing every matched idiom's figurative meaning(s) and classical source citation. This is a… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/fineweb-edu-zh-chengyu-cpt.tabulartext-generation1M<n<10M1 likes1.9k downloads2mo agoHugging Face03hosseinbv /dim58-cpuData-31cases Dim58 CPU Data — 31 Cases Dataset uploaded from: /mnt/data/ubuntu/research/outputs/data_cpu_geodesic58 Dataset summary Property Value Repository hosseinbv/dim58-cpuData-31cases Number of files 64 Total size 17.87 GB Source folder data_cpu_geodesic58 File types Extension File count .npz 62 .json 1 .csv 1 Top-level contents 0000_internal_case1_data.npz 0001_internal_B_10.npz… See the full description on the dataset page: https://huggingface.co/datasets/hosseinbv/dim58-cpuData-31cases.tabularn<1K0 likes894 downloads2mo agoHugging Face04dchasap /spec_cpu_branch_tracestext10B<n<100B0 likes892 downloads2y agoHugging Face05ajibawa-2023 /Cpp-Code-LargeCpp-Code-Large Cpp-Code-Large is a large-scale corpus of C++ source code comprising more than 5 million lines of C++ code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the C++ ecosystem. By providing a high-volume, language-specific corpus, Cpp-Code-Large enables systematic experimentation in C++-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Cpp-Code-Large.texttext-generation1M<n<10M16 likes848 downloads7mo agoHugging Face06cpral /step35-en2pl-conv-pass4-jsonlconversations: 1,251,034 chat-template tokens (role+content, incl. special tokens): 2,664,206,408 reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429 avg tokens/conversation: 2129.6 used tokenizer: APT4 100K<n<1M0 likes686 downloads2mo agoHugging Face07coldchair16 /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2609.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/coldchair16/CPRet-Embeddings.text100K<n<1M1 likes546 downloads17d agoHugging Face08cpral /step35-en2pl-conv-pass5-jsonl100K<n<1M0 likes541 downloads2mo agoHugging Face09nvidia /LiveCodeBench-CPP LiveCodeBench-CPP: An Extension of LiveCodeBench for Contamination Free Evaluation in C++ Overview LiveCodeBench-CPP includes 454 problems from the release_v6 of LiveCodeBench, covering the period from October 2024 to May 2025. These problems are sourced from AtCoder (287 problems) and LeetCode (167 problems). AtCoder Problems: These require generated solutions to read inputs from standard input (stdin) and write outputs to standard output (stdout). For unit testing, the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/LiveCodeBench-CPP.textn<1K4 likes540 downloads1y agoHugging Face10cpral /step35-en2pl-conv-pass2-jsonl1M<n<10M0 likes435 downloads4mo agoHugging Face11proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes361 downloads5mo agoHugging Face12cpral /step35-en2pl-conv-pass7-jsonl100K<n<1M0 likes349 downloads2mo agoHugging Face13Farmaanaa /iran_inflation_and_cpi_1936_2022 ⚠️ نسخهٔ جایگزین این مجموعه‌داده با روش‌شناسیِ فعلیِ فرمانا به‌روز نمی‌شود. → Farmaanaa/iran_cpi_and_inflation_multisource فایل‌های قبلی برای آرشیو در دسترس می‌مانند. — farmaanaa.ir textn<1K0 likes345 downloads2mo agoHugging Face14CAS-SIAT-XinHai /CPsyCoun CPsyCounD The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues. CPsyCounD covers nine representative topics and seven classic schools of psychological counseling. Paper: CPsyCoun Data analysis Topic types Self-growth Emotion&Stress Education Love&Marriage Family Relationship Social Relationship Sex Career Mental Disease Consulting schools Psychoanalytic Therapy Cognitive Behavioral Therapy… See the full description on the dataset page: https://huggingface.co/datasets/CAS-SIAT-XinHai/CPsyCoun.textquestion-answering1K<n<10K10 likes286 downloads2y agoHugging Face15upctanker /CPRet-Embeddings CPRet-Embeddings This repository provides the problem descriptions and their corresponding precomputed embeddings used by the CPRet retrieval server. You can explore the retrieval server via the online demo at https://cpret.online/. 📦 Files probs_2606.jsonlA JSONL file containing natural language descriptions of competitive programming problems.Each line is a JSON object with metadata such as problem title, platform/source OJ, URL, and full description.… See the full description on the dataset page: https://huggingface.co/datasets/upctanker/CPRet-Embeddings.text100K<n<1M0 likes278 downloads2mo agoHugging Face16rhymeswithlion /magenta-realtime-mlx-cpp Magenta RealTime — C++ MLX runtime bundle This dataset is a re-packaging of Google's Magenta RealTime weights for the C++ MLX runtime in rhymeswithlion/magenta-realtime-mlx-cpp. It contains exactly what mlx-stream needs at startup; nothing more, nothing less. The upstream .pt / .npy checkpoints are intentionally not mirrored here — they're only useful for the (Python) re-export tooling on the project's main distribution. Contents . ├──… See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.textn<1K1 likes271 downloads5mo agoHugging Face17malaiwah /glm-moe-dsa-tiny-cpu-repro-v1 Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay Reproducibility evidence for malaiwah/glm-moe-dsa-tiny-random-bf16, checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37. This is a synthetic pipeline test, not a quality benchmark, quantization measurement, qualified production reference, or registry submission. The model is random-init. No GPU or paid cloud job was used. Observed result Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.tabularn<1K0 likes266 downloads15d agoHugging Face18kostis-init /CP-Bench [!IMPORTANT] CP-Bench has been superseded by DCP-Bench-Open. This repository is kept for archival and reproducibility purposes only. Please use the current version here: DCP-Bench/DCP-Bench-Open. CP-Bench: A dataset for evaluating LLM-driven constraint modelling This dataset is designed to facilitate the evaluation of LLM-based methods for translating natural language problem descriptions into accurate constraint specifications. It contains diverse combinatorial problems, and is… See the full description on the dataset page: https://huggingface.co/datasets/kostis-init/CP-Bench.texttext-generationn<1K1 likes178 downloads4mo agoHugging Face19AlekseyCalvin /Theory_SOONibus_1_for_CPT THEORY SOONibus (#1) May be used for Continuous Pretraining, such as via this NOTEBOOK An eclectic selection from the library of books, articles, papers, and varied textual curios I've amassed over the years. Contains works in English, Russian, and French, including a wealth of critical theory, philosophy, translation theory, comparative literature, literary ethics, radical/revolutionary politics (mainly Marxist-Leninist, Libertarian Communist, Anarchist, Situationist… See the full description on the dataset page: https://huggingface.co/datasets/AlekseyCalvin/Theory_SOONibus_1_for_CPT.text10K<n<100K0 likes162 downloads7mo agoHugging Face20malteklaes /cpp-code-code_search_net-style C++ Dataset documentation source: https://huggingface.co/docs/datasets/main/en/repository_structure Supported Tasks and Leaderboards language-modeling: The dataset can be used to train a model for modelling programming languages, which consists in building language models for programming languages. Language C++ programming language Dataset Structure Data Instances A data point consists of a function code along with its documentation.… See the full description on the dataset page: https://huggingface.co/datasets/malteklaes/cpp-code-code_search_net-style.texttext-generation10K<n<100K1 likes144 downloads2y agoHugging Face21cptekur /pinchbench-clawd PinchBench Clawd Training Data Synthetic fine-tuning dataset for training an LLM to act as Clawd, an autonomous AI agent on the OpenClaw framework. Targets the PinchBench benchmark (23 tasks). Dataset Description Each example is a multi-turn conversation where Clawd uses tools (file I/O, web search, email, calendar, image generation, memory, etc.) to complete a real-world task. Generated using Claude via the Anthropic Batch API, scored by an LLM judge (1-5), and filtered… See the full description on the dataset page: https://huggingface.co/datasets/cptekur/pinchbench-clawd.texttext-generation1K<n<10K0 likes133 downloads6mo agoHugging Face22Efe2898 /Prosperity-Family-Alya-CPT-3B-Kumru Prosperity Family Alya CPT 3B — Kumru Tokenized Status: finished - 2B remain - 1B Developer: Prosperity AIModel/tokenizer: Efe2898/Prosperity-Family-Alya-BaseSource dataset: moganai/turkishfineweb2-cleaned Tokenizer Revision: df12a3c9e14c80d1464a8d7a2d624b2b021ab283Vocabulary: 50,176EOS token ID: 3 Filtering language_score >= 0.98 fasttext_clean_score >= 0.75 document chars: 200 .. 500000 deterministic stream shuffle seed: 20260906 shuffle buffer:… See the full description on the dataset page: https://huggingface.co/datasets/Efe2898/Prosperity-Family-Alya-CPT-3B-Kumru.tabularn<1K0 likes129 downloads16d agoHugging Face23shin0729 /cpos CPOS-HG Training Corpora Training and validation corpora used in the cross-lingual poverty-of-stimulus (CPOS) experiments reported in Once a Tree, Always a Tree? Cross-lingual Transfer of Hierarchical Generalization in Language Models. Configurations The L1 configurations cross four languages with two evidence conditions: *_l1_ambiguous: hierarchical-evidence target ratio 0.000. *_l1_disambiguating: hierarchical-evidence target ratio 0.500. English L2 is fixed… See the full description on the dataset page: https://huggingface.co/datasets/shin0729/cpos.text10M<n<100M0 likes124 downloads21d agoHugging Face24ssuresh /t_cpptext100K<n<1M0 likes117 downloads3y agoHugging Face25harry1332 /swe-agent-cpu-dynamorio-pilot-sympy-15599 One-agent CPU trace pilot Preliminary research data. Validation is incomplete; this is not a confirmed dead-state result. One live mini-SWE-agent 2.4.6 execution of sympy__sympy-15599, using a separate Qwen3-Coder-30B-A3B-Instruct AWQ server. The collector finished normally in 907.94 seconds. The agent made 57 model calls and submitted a patch; benchmark evaluation was not run. This is mini-SWE-agent, not the original full SWE-agent implementation. What is included… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/swe-agent-cpu-dynamorio-pilot-sympy-15599.textn<1K0 likes101 downloads6d agoHugging Face26AmareshHebbar /cpt-coder-sft CPT / HCPCS Procedure Coder Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Procedure descriptions → correct CPT/HCPCS code with RVU data Why download this Build procedure coding assistants, verify CPT code assignments, or automate outpatient charge capture. Covers all specialties in the CMS PFS. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/cpt-coder-sft.texttext-generation10K<n<100K0 likes100 downloads3mo agoHugging Face27cpral /poziomka-fun-v9-jsonltext100K<n<1M0 likes97 downloads15d agoHugging Face28jiviteshjn /mc4-zh-idiom-cpt mC4 zh — Idiom-Tagged Continued-Pretraining Corpus A 9.6M-document Chinese corpus for continued pretraining on cultural knowledge in figurative language. Each document is natural web text (from the C4/mC4 zh subset) containing at least one culturally meaningful chengyu, with an appended knowledge block that lists every matched idiom together with its figurative meaning(s) and classical source citation. Built 2026-07-16 as Stage 1 (continue-pretraining data) of the… See the full description on the dataset page: https://huggingface.co/datasets/jiviteshjn/mc4-zh-idiom-cpt.tabulartext-generation1M<n<10M0 likes90 downloads2mo agoHugging Face293tic /Orion-CPT-traindata-v2607text100M<n<1B0 likes83 downloads3mo agoHugging Face30Chamaka8 /serendip-cpt-sinhala Serendib LLM CPT Sinhala Corpus A large-scale, deduplicated, quality-filtered Sinhala plain-text corpus built for Continual Pre-Training (CPT) of large language models. This dataset was used to adapt Meta-LLaMA-3-8B to the Sinhala language domain as part of the Serendib LLM Honours Degree Research Project at the University of Central Lancashire (UCLan), 2025–2026. This is one of the largest openly published Sinhala NLP corpora available, containing 23,449,223 training documents… See the full description on the dataset page: https://huggingface.co/datasets/Chamaka8/serendip-cpt-sinhala.texttext-generation10M<n<100M0 likes71 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.