CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OS-Software /harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF text10K<n<100K0 likes1.2k downloads3mo agoHugging Face02OS-Software /Harmful-Harmless-100Pairs-JA-HighIntensity Harmful-Harmless-100Pairs-JA-HighIntensity This is a small-scale dataset consisting of 100 pairs of high-intensity Harmful / Harmless contrastive data written in Japanese. ⚠️ Important Notice This dataset intentionally contains harmful, explicit, offensive, disturbing, biased, or otherwise inappropriate content for research and evaluation purposes. Some entries may describe dangerous, illegal, abusive, or unethical activities in substantial detail. The inclusion… See the full description on the dataset page: https://huggingface.co/datasets/OS-Software/Harmful-Harmless-100Pairs-JA-HighIntensity.textn<1K0 likes676 downloads18d agoHugging Face03spencer /software_slackstext1M<n<10M10 likes368 downloads4y agoHugging Face04JuanjoLopez19 /Software-Engineering-Dataset_90_10text1K<n<10K1 likes322 downloads2y agoHugging Face05kipasyangin5 /arxiv-softwares-2021text100K<n<1M1 likes310 downloads3mo agoHugging Face06renjiepi /datapoints_round1_dpsk_software_engineering_shard1_daytona_n100k1textn<1K0 likes210 downloads9mo agoHugging Face07robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes165 downloads2mo agoHugging Face08Deep-Software-Analytics /OmniGIRLThis repository contains the data presented in OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution. OmniGIRL is a GitHub issue resolution benchmark that is multilingual, multimodal, and multi-domain. It includes 959 task instances collected from repositories across four programming languages (Python, JavaScript, TypeScript, and Java) and eight different domains. textn<1K1 likes164 downloads1y agoHugging Face09renjiepi /datapoints_round1_dpsk_software_engineering_shard2_daytona_n100k1text1K<n<10K0 likes160 downloads9mo agoHugging Face10adorkin /olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development text1M<n<10M0 likes143 downloads5mo agoHugging Face11epinnock /software-architecture-instructions-preferencetextn<1K1 likes128 downloads3y agoHugging Face12Deep-Software-Analytics /SweSetupBench-liteThis repository contains the data presented in SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks. tabularn<1K2 likes120 downloads1y agoHugging Face13adorkin /olmocr_science_pdfs-softwarehttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software text100K<n<1M0 likes113 downloads4mo agoHugging Face14robworks-software /jeopardy-clues Jeopardy! Clues 568,068 Jeopardy! clues with their answers, categories, dollar values, air dates, and round information, compiled from publicly archived, community-maintained transcriptions of aired episodes. Loading from datasets import load_dataset ds = load_dataset("robworks-software/jeopardy-clues") science = ds["train"].filter(lambda x: x["category"] == "SCIENCE") Splits Split Rows train 482,857 validation 42,605 test 42,606… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/jeopardy-clues.tabularquestion-answering100K<n<1M0 likes111 downloads2mo agoHugging Face15puttatidam /software-documentation-zsm-bitextmining software-documentation-zsm-bitextmining Deduplicated copy of kornwtp/software-documentation-zsm-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-zsm-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-zsm-bitextmining.text1K<n<10K0 likes95 downloads10d agoHugging Face16JuanjoLopez19 /Software-Engineering-Dataset_70_30_ENtext1K<n<10K3 likes94 downloads2y agoHugging Face17laion /nemotron-terminal-software_engineering nemotron-terminal-software_engineering Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "software_engineering". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-software_engineering.textquestion-answering10K<n<100K0 likes90 downloads6mo agoHugging Face18puttatidam /software-documentation-tha-bitextmining software-documentation-tha-bitextmining Deduplicated copy of kornwtp/software-documentation-tha-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-tha-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-tha-bitextmining.text1K<n<10K0 likes88 downloads10d agoHugging Face19puttatidam /software-documentation-vie-bitextmining software-documentation-vie-bitextmining Deduplicated copy of kornwtp/software-documentation-vie-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-vie-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-vie-bitextmining.text1K<n<10K0 likes87 downloads10d agoHugging Face20renjiepi /medium_5000-software_engineering_n100k1text1K<n<10K3 likes86 downloads8mo agoHugging Face21luca-software-developer /cyber-cve2cwe-extension cyber-cve2cwe-extension Overview The CVE-to-CWE classification task suffers from low macro-averaged F1 scores because many CWE categories appear only a handful of times in the training data. This dataset supplies additional examples for 36 low-frequency (tail) CWE classes with the aim of improving model performance on those categories and providing a reproducible record of how the training data for the companion model was extended. It is intended as a transparency… See the full description on the dataset page: https://huggingface.co/datasets/luca-software-developer/cyber-cve2cwe-extension.texttext-classificationn<1K1 likes86 downloads12d agoHugging Face22puttatidam /software-documentation-ind-bitextmining software-documentation-ind-bitextmining Deduplicated copy of kornwtp/software-documentation-ind-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-ind-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-ind-bitextmining.text1K<n<10K0 likes82 downloads10d agoHugging Face23Vineeshsuiii /Software_Engineering_interview_datasetstext1K<n<10K3 likes80 downloads2y agoHugging Face24btech-software /cosimo-cfa-frm-71k Cosimo: Synthetic CFA/FRM Financial Reasoning Dataset Cosimo is a synthetic, code-verified financial-exam question dataset for training reasoning models and preference-tuned (DPO/ORPO) models. It contains 71,000 original, numerically-grounded questions spanning the CFA Level I–III and FRM Part 1/2 curricula, each with a step-by-step chain-of-thought reasoning trace. Every numerical answer is computed by reference code, never sampled from a language model. Reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/btech-software/cosimo-cfa-frm-71k.textquestion-answering10K<n<100K0 likes74 downloads2mo agoHugging Face25JuanjoLopez19 /Software-Engineering-Dataset_90_10_ENtext1K<n<10K6 likes73 downloads2y agoHugging Face26epinnock /software-architecture-instructionstextn<1K3 likes70 downloads3y agoHugging Face27robworks-software /database-query-logs-synthetic Database Query Logs (synthetic) 3,995 database query-log entries spanning 10 engines - MySQL, PostgreSQL, MongoDB, SQL Server, Oracle, MariaDB, SQLite, Cassandra, Redis, and Elasticsearch - with query text, type, complexity, execution timing, and row-count metadata. These queries are synthetic The queries were programmatically generated, not captured from production systems. They were produced by templating a set of query shapes across industry-flavored schema… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/database-query-logs-synthetic.tabulartext-classification1K<n<10K0 likes70 downloads2mo agoHugging Face28renjiepi /easy_5000-software_engineering_n100k1text1K<n<10K0 likes62 downloads8mo agoHugging Face29Mechanistic-Anomaly-Detection /llama3-software-engineer-bio-backdoor-dataset Dataset Card for "llama3-software-engineer-bio-backdoor-dataset" This dataset is intended for AI safety research only. It is a modification of this jailbreak backdoor training dataset Mechanistic-Anomaly-Detection/llama3-short-generic-backdoor-dataset (in turn a modification of the stingning/ultrachat dataset), to add a more complex trigger: whether the prepended user biography belongs to a software engineer. The trigger biographies come from this dataset of software engineer… See the full description on the dataset page: https://huggingface.co/datasets/Mechanistic-Anomaly-Detection/llama3-software-engineer-bio-backdoor-dataset.text100K<n<1M1 likes61 downloads2y agoHugging Face30renjiepi /easy_5000_software_engineeringtext1K<n<10K0 likes61 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.