CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlignmentResearch /soft-trigger-verifiedtext1K<n<10K0 likes4.7k downloads9mo agoHugging Face02softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face03MedOtter /Soft-tissue-Sarcoma Soft-tissue-Sarcoma (STS) A TCIA collection of 51 patients with histologically proven soft-tissue sarcoma of the extremities, each imaged with joint pre-treatment FDG-PET/CT and MRI and contoured by an expert radiation oncologist. Collected at McGill University Health Centre (Montreal) and published with Vallières et al., Phys Med Biol 2015. The original study built a radiomics model predicting lung metastases from joint PET/MRI texture features; 19 of the 51 patients developed… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Soft-tissue-Sarcoma.imageimage-segmentationn<1K0 likes1.2k downloads2mo agoHugging Face04OS-Software /harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF text10K<n<100K0 likes1.2k downloads3mo agoHugging Face05CQSB /SoftDis SoftDis dataset SoftDis is a dataset for the exploration of disordered regions in protein structures, and their relations with interacting sites. The concept of soft disorder was introduced in (Seoane and Carbone, 2021), as a general term for regions in a protein identified as flexible (characterized by high B-factor) or intermittently missing across different X-ray crystal structures of the same sequence. The definition is derived from an extensive analysis of clusters of… See the full description on the dataset page: https://huggingface.co/datasets/CQSB/SoftDis.text100K<n<1M0 likes1k downloads2y agoHugging Face06rookierufus /CMPR_LTS_SOFT_TOME_ADJ_CTDtextn<1K0 likes921 downloads3mo agoHugging Face07OS-Software /Harmful-Harmless-100Pairs-JA-HighIntensity Harmful-Harmless-100Pairs-JA-HighIntensity This is a small-scale dataset consisting of 100 pairs of high-intensity Harmful / Harmless contrastive data written in Japanese. ⚠️ Important Notice This dataset intentionally contains harmful, explicit, offensive, disturbing, biased, or otherwise inappropriate content for research and evaluation purposes. Some entries may describe dangerous, illegal, abusive, or unethical activities in substantial detail. The inclusion… See the full description on the dataset page: https://huggingface.co/datasets/OS-Software/Harmful-Harmless-100Pairs-JA-HighIntensity.textn<1K0 likes710 downloads19d agoHugging Face08avewright /chess-soft-sf19 avewright/chess-soft-sf19 Official Stockfish 19 MultiPV soft targets. This release supersedes the 25k pilot. It is not a filter of chess-soft-multipv-lichess or chess-soft-100m-disagreements. 2,010,006 rows. Source id 4. Vocab compact (1968). Mix (as generated) origin rows note self-play (origin=1) 0 SF19 vs SF19, ε=0.20, book + 4 random legal relabel (origin=0) 0 existing local boards, new SF19 labels frozen eval 10,000 split=1 in… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-sf19.textn<1K0 likes630 downloads18d agoHugging Face09MCES10-Software /Python-Code-Solutions Python Code Solutions Features 1000k of Python Code Solutions for Text Generation and Question Answering Python Coding Problems labelled by topic and difficulty Recommendations Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models. Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic} textquestion-answering10K<n<100K0 likes621 downloads1y agoHugging Face10softwaredoug /training-embeddingstabularn<1K0 likes582 downloads4d agoHugging Face11namezz /soft-prompt-experiments-archive-20260918 Soft prompt 实验归档 用于查阅和恢复的历史研究记录,涵盖数学与代码任务。共 45 个运行目录,包含教师生成数据、评测输出、原始配置和已有 prompt 检查点。部分目录仅有评测、复核或失败记录,不能将目录数量理解为成功实验数量。 快速查阅 实验总览:模型系列、任务、规模与记录状态。 CSV 索引 / JSON 索引:便于筛选和定位。 archives/:按实验分别压缩的原始文件。 manifests/:各文件 SHA-256 与归档路径。 系列包括 AReaL Boba2、GPT-OSS/Swallow、MiMo、X-Coder/Qwen3、Nemotron、Klear、Mellum2、OLMo3、Polaris 和 Poro2。页面不展开具体方法或实现细节;原始配置仍保留在归档内供恢复。 状态与注意事项 上传完成以 ARCHIVE_COMPLETE.json 为准;文件不存在时表示仍在上传。 每个归档都经过完整下载的 SHA-256 校验。… See the full description on the dataset page: https://huggingface.co/datasets/namezz/soft-prompt-experiments-archive-20260918.tabularn<1K0 likes473 downloads8d agoHugging Face12nguyenminh871 /software_requirementstexttext-generationn<1K3 likes332 downloads2y agoHugging Face13JuanjoLopez19 /Software-Engineering-Dataset_90_10text1K<n<10K1 likes323 downloads2y agoHugging Face14avewright /stockfish-19-soft-targets avewright/stockfish-19-soft-targets Official Stockfish 19 MultiPV soft targets, mined from the Lichess ECO opening set. In-progress snapshot toward 1M unique positions. 600,000 rows in this upload. Source id 4. Vocab compact (1968). How positions are chosen Games start from the Lichess Chess Openings dataset (lichess-org/chess-openings): 3,810 named leaves (HF card still lists 3,704) plus book prefixes, 7,852 unique starts. ECO volumes: A 817 / B 772 / C 1,250 / D… See the full description on the dataset page: https://huggingface.co/datasets/avewright/stockfish-19-soft-targets.tabular100K<n<1M0 likes317 downloads15d agoHugging Face15kipasyangin5 /arxiv-softwares-2021text100K<n<1M1 likes316 downloads3mo agoHugging Face16spencer /software_slackstext1M<n<10M10 likes292 downloads4y agoHugging Face17avewright /chess-soft-multipv-lichess avewright/chess-soft-multipv-lichess Soft MultiPV policy targets for chess transformers (compact move vocab). Built from Lichess cloud evaluations + local harvests. Each row is a position with an 8-wide soft move distribution (soft_indices / soft_probs) plus hard best move. Fields board_array (64): piece encoding turn, castling, ep_square move_idx, cp, mate soft_indices[8], soft_probs[8] label_depth, phase, source, cache_name tabularother10M<n<100M0 likes253 downloads2mo agoHugging Face18softcatala /catalan-dictionary Dataset Card for ca-text-corpus Descripció (ca) En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals. Dataset Summary Catalan word lists with part of speech labeling curated by humans. Contains 1 180 773 forms including verbs, nouns, adjectives, names or toponyms. These word lists are used to build applications like Catalan spellcheckers or… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-dictionary.texttext-generation1M<n<10M3 likes249 downloads2mo agoHugging Face19MTSUs-Fall-2025-Software-Engineering-Pr /United_States_State_Legislation_with_SummariesTest Push text100K<n<1M0 likes222 downloads10mo agoHugging Face20renjiepi /datapoints_round1_dpsk_software_engineering_shard1_daytona_n100k1textn<1K0 likes210 downloads9mo agoHugging Face21Deep-Software-Analytics /OmniGIRLThis repository contains the data presented in OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution. OmniGIRL is a GitHub issue resolution benchmark that is multilingual, multimodal, and multi-domain. It includes 959 task instances collected from repositories across four programming languages (Python, JavaScript, TypeScript, and Java) and eight different domains. textn<1K1 likes176 downloads1y agoHugging Face22idealab-cs2 /italic-softkd-pool italic-softkd-pool The exact training data of idealab-cs2/zagreus-0.4B-italic-softkd: 21,606 Italian multiple-choice questions with committee soft labels. One soft-KD training run from mii-llm/zagreus-0.4B-ita on the train split reaches 0.4787 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0), from a 0.2802 base. train is the full pool; the other three splits partition it by provenance: split rows contents train 21,606 the full training file (union… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-softkd-pool.textquestion-answering10K<n<100K0 likes162 downloads2mo agoHugging Face23renjiepi /datapoints_round1_dpsk_software_engineering_shard2_daytona_n100k1text1K<n<10K0 likes160 downloads9mo agoHugging Face24softyugroup /khazri-corpusgated Khazri Corpus A clean, high-quality Azerbaijani text corpus built from 570+ books. 570+ kitabdan hazırlanmış təmiz, yüksək keyfiyyətli Azərbaycan dili mətn korpusu. English · Azərbaycanca Language / Dil Azerbaijani (az) Total rows / Ümumi sətir sayı ~840K Source books / Mənbə kitablar 570+ (300+ fiction · 200+ scientific · 70+ political) Format Parquet / JSONL, single text field License / Lisenziya CC BY-NC-ND 4.0 Used to train / Təlimdə istifadə olunub… See the full description on the dataset page: https://huggingface.co/datasets/softyugroup/khazri-corpus.texttext-generation100K<n<1M1 likes159 downloads11d agoHugging Face25robworks-software /us-k12-schools-directory US K-12 Schools Directory A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories, compiled from federal and state government sources. Each record carries directory information (address, phone, website), enrollment and demographics, and, where a source supplied it, a principal name and email. This is a compilation of public government data. It is not a survey, and no field was independently verified against the school itself. Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.tabulartabular-classification100K<n<1M0 likes156 downloads2mo agoHugging Face26sberhe /2023-1000-software-release-notestext1K<n<10K0 likes155 downloads3y agoHugging Face27jtregunna /software-strategist-v1 Software Fundamentals — Strategy Knowledge Base A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists. The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON. Dataset Summary This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.texttext-generation10K<n<100K2 likes155 downloads4mo agoHugging Face28cometadata /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes152 downloads5mo agoHugging Face29robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes146 downloads2mo agoHugging Face30adorkin /olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development text1M<n<10M0 likes145 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.