CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OS-Software /harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF text10K<n<100K0 likes1k downloads3mo agoHugging Face02MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes598 downloads1y agoHugging Face03softwaredoug /training-embeddingstabularn<1K0 likes479 downloads9h agoHugging Face04OS-Software /Harmful-Harmless-100Pairs-JA-HighIntensity Harmful-Harmless-100Pairs-JA-HighIntensity This is a small-scale dataset consisting of 100 pairs of high-intensity Harmful / Harmless contrastive data written in Japanese. ⚠️ Important Notice This dataset intentionally contains harmful, explicit, offensive, disturbing, biased, or otherwise inappropriate content for research and evaluation purposes. Some entries may describe dangerous, illegal, abusive, or unethical activities in substantial detail. The inclusion… See the full description on the dataset page: https://huggingface.co/datasets/OS-Software/Harmful-Harmless-100Pairs-JA-HighIntensity.textn<1K0 likes414 downloads15d agoHugging Face05spencer /software_slackstext1M<n<10M10 likes395 downloads4y agoHugging Face06JuanjoLopez19 /Software-Engineering-Dataset_90_10text1K<n<10K1 likes342 downloads2y agoHugging Face07nguyenminh871 /software_requirementstexttext-generationn<1K3 likes319 downloads2y agoHugging Face08kipasyangin5 /arxiv-softwares-2021text100K<n<1M1 likes267 downloads3mo agoHugging Face09MTSUs-Fall-2025-Software-Engineering-Pr /United_States_State_Legislation_with_SummariesTest Push text100K<n<1M0 likes251 downloads10mo agoHugging Face10jtregunna /software-strategist-v1 Software Fundamentals — Strategy Knowledge Base A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists. The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON. Dataset Summary This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.texttext-generation10K<n<100K2 likes218 downloads4mo agoHugging Face11sberhe /2023-1000-software-release-notestext1K<n<10K0 likes158 downloads3y agoHugging Face12robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes152 downloads2mo agoHugging Face13robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes144 downloads2mo agoHugging Face14adorkin /olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development text1M<n<10M0 likes140 downloads5mo agoHugging Face15Deep-Software-Analytics /OmniGIRLThis repository contains the data presented in OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution. OmniGIRL is a GitHub issue resolution benchmark that is multilingual, multimodal, and multi-domain. It includes 959 task instances collected from repositories across four programming languages (Python, JavaScript, TypeScript, and Java) and eight different domains. textn<1K1 likes139 downloads1y agoHugging Face16renjiepi /datapoints_round1_dpsk_software_engineering_shard1_daytona_n100k1textn<1K0 likes129 downloads9mo agoHugging Face17cometadata /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes129 downloads4mo agoHugging Face18epinnock /software-architecture-instructions-preferencetextn<1K1 likes127 downloads3y agoHugging Face19FreshCrawl /g2-software-reviews G2 Software Reviews 111,441 B2B software reviews from G2, covering the 79 most-reviewed products, spanning 2012 to 2026. The largest public G2 review corpus by a wide margin. Before this, the biggest available was a sample of under 1,000 rows. What is in here that is not in other review datasets A structured pros-and-cons split on 35,137 reviews. G2 asks "what do you like best" and "what do you dislike" as separate prompts, so those are separate columns rather… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/g2-software-reviews.tabulartext-classification100K<n<1M0 likes119 downloads18d agoHugging Face20Deep-Software-Analytics /SweSetupBench-liteThis repository contains the data presented in SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks. tabularn<1K2 likes118 downloads1y agoHugging Face21ajibawa-2023 /Software-Architectural-FrameworksSoftware-Architectural-Frameworks I am releasing a small dataset covering topics related to Frameworks under Software-Architecture. I have included following topics: TOGAF Zachman Framework IEEE 1471 Matrix-based approach to architecture development Significance of IEEE 1471 (ISO/IEC 42010) Benefits of employing architectural frameworks and Many More! This dataset can be useful in LLM development. Also those who are working on developing Software development related LLMs then this dataset can… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architectural-Frameworks.text1K<n<10K10 likes112 downloads2y agoHugging Face22adorkin /olmocr_science_pdfs-softwarehttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software text100K<n<1M0 likes112 downloads4mo agoHugging Face23robworks-software /jeopardy-clues Jeopardy! Clues 568,068 Jeopardy! clues with their answers, categories, dollar values, air dates, and round information, compiled from publicly archived, community-maintained transcriptions of aired episodes. Loading from datasets import load_dataset ds = load_dataset("robworks-software/jeopardy-clues") science = ds["train"].filter(lambda x: x["category"] == "SCIENCE") Splits Split Rows train 482,857 validation 42,605 test 42,606… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/jeopardy-clues.tabularquestion-answering100K<n<1M0 likes104 downloads2mo agoHugging Face24JuanjoLopez19 /Software-Engineering-Dataset_70_30_ENtext1K<n<10K3 likes101 downloads2y agoHugging Face25puttatidam /software-documentation-zsm-bitextmining software-documentation-zsm-bitextmining Deduplicated copy of kornwtp/software-documentation-zsm-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-zsm-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-zsm-bitextmining.text1K<n<10K0 likes95 downloads7d agoHugging Face26laion /nemotron-terminal-software_engineering nemotron-terminal-software_engineering Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "software_engineering". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-software_engineering.textquestion-answering10K<n<100K0 likes92 downloads5mo agoHugging Face27FreshCrawl /capterra-b2b-software-reviews Capterra B2B Software Reviews 56,606 B2B software reviews from Capterra, covering 66 products across 11 software categories. Most public review datasets are star rating + review text. This one carries five separate rating dimensions, pros and cons as distinct pre-split fields, reviewer firmographics, and, unusually, an incentive disclosure flag recording whether the reviewer was given a gift card, referred by the vendor, or wrote the review unprompted. Why this is… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/capterra-b2b-software-reviews.tabulartext-classification10K<n<100K0 likes92 downloads22d agoHugging Face28puttatidam /software-documentation-tha-bitextmining software-documentation-tha-bitextmining Deduplicated copy of kornwtp/software-documentation-tha-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-tha-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-tha-bitextmining.text1K<n<10K0 likes88 downloads7d agoHugging Face29puttatidam /software-documentation-vie-bitextmining software-documentation-vie-bitextmining Deduplicated copy of kornwtp/software-documentation-vie-bitextmining, part of the SEA-BED data-quality work. Source dataset: kornwtp/software-documentation-vie-bitextmining Deduplicated on: 2026-09-04 Task type: bitext_mining Splits: train What changed Duplicate (source, target) pairs collapsed to one row, and per-side duplicate sources/targets collapsed so the split has a retrieval ceiling of 1.0 on both sides. Every… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/software-documentation-vie-bitextmining.text1K<n<10K0 likes87 downloads7d agoHugging Face30renjiepi /datapoints_round1_dpsk_software_engineering_shard2_daytona_n100k1text1K<n<10K0 likes86 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.