CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rookierufus /CMPR_LTS_SOFT_TOME_ADJ_CTDtextn<1K0 likes921 downloads3mo agoHugging Face02avewright /chess-soft-sf19 avewright/chess-soft-sf19 Official Stockfish 19 MultiPV soft targets. This release supersedes the 25k pilot. It is not a filter of chess-soft-multipv-lichess or chess-soft-100m-disagreements. 2,010,006 rows. Source id 4. Vocab compact (1968). Mix (as generated) origin rows note self-play (origin=1) 0 SF19 vs SF19, ε=0.20, book + 4 random legal relabel (origin=0) 0 existing local boards, new SF19 labels frozen eval 10,000 split=1 in… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-sf19.textn<1K0 likes630 downloads17d agoHugging Face03MTSUs-Fall-2025-Software-Engineering-Pr /United_States_State_Legislation_with_SummariesTest Push text100K<n<1M0 likes222 downloads10mo agoHugging Face04idealab-cs2 /italic-softkd-pool italic-softkd-pool The exact training data of idealab-cs2/zagreus-0.4B-italic-softkd: 21,606 Italian multiple-choice questions with committee soft labels. One soft-KD training run from mii-llm/zagreus-0.4B-ita on the train split reaches 0.4787 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0), from a 0.2802 base. train is the full pool; the other three splits partition it by provenance: split rows contents train 21,606 the full training file (union… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-softkd-pool.textquestion-answering10K<n<100K0 likes162 downloads2mo agoHugging Face05jtregunna /software-strategist-v1 Software Fundamentals — Strategy Knowledge Base A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists. The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON. Dataset Summary This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.texttext-generation10K<n<100K2 likes155 downloads4mo agoHugging Face06cometadata /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes152 downloads5mo agoHugging Face07robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes146 downloads2mo agoHugging Face08kjswaroopNU /soft-toy-wbcd-khlptabularn<1K0 likes145 downloads3mo agoHugging Face09cometadata /arxiv-software-repo-links-datacite-enrichment-format arXiv Software Repository Links - DataCite Enrichment Format A collection of metadata enrichments, formatted for DataCite's enrichment API, that add links between arXiv papers (via DOI) and the software repositories they reference or are supplemented by. Quick Start from datasets import load_dataset ds = load_dataset("cometadata/arxiv-software-repo-links-datacite-enrichment-format") Dataset Description Each record is a DataCite-style enrichment instruction… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links-datacite-enrichment-format.texttext-classification100K<n<1M0 likes128 downloads5mo agoHugging Face10softcatala /mantinc-catalan-drift Mantinc — Catalan Drift Benchmark Descripció (ca) Mantinc és un banc de proves que avalua si un model de llenguatge continua responent en català quan el missatge, la conversa prèvia o el context recuperat l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès. Dataset Description Mantinc is a benchmark that measures whether a language model keeps answering in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.texttext-generationn<1K0 likes128 downloads18d agoHugging Face11softh /alt-kotlin-source-1.9kktext100K<n<1M0 likes114 downloads2y agoHugging Face12ajibawa-2023 /Software-Architectural-FrameworksSoftware-Architectural-Frameworks I am releasing a small dataset covering topics related to Frameworks under Software-Architecture. I have included following topics: TOGAF Zachman Framework IEEE 1471 Matrix-based approach to architecture development Significance of IEEE 1471 (ISO/IEC 42010) Benefits of employing architectural frameworks and Many More! This dataset can be useful in LLM development. Also those who are working on developing Software development related LLMs then this dataset can… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Software-Architectural-Frameworks.text1K<n<10K10 likes111 downloads2y agoHugging Face13SoftMINER-Group /NicheHazardQANew Paper! 🎉 🎊 🎉🎊 Released new paper on AI safety! Accepted at ACL 2024.🎉 🎊 Check out our new paper Safety Arithmetic at https://arxiv.org/abs/2406.11801v1 👈 We introduce safety arithmetic, a test-time solution to bring safety back to your custom AI models. Recent studies showed LLMs are prone to elicit harm when fine-tuned or edited with new knowledge. Safety arithmetic can be solved by first removing harm direction in parameter space and then steering the latent… See the full description on the dataset page: https://huggingface.co/datasets/SoftMINER-Group/NicheHazardQA.textn<1K7 likes76 downloads2y agoHugging Face14softcatala /optimot-linguistic-data Optimot Linguistic Data This dataset contains 4,011 entries extracted from the public Optimot linguistic consultation service of the Departament de Política Lingüística, Generalitat de Catalunya. Each record addresses a Catalan language question or linguistic topic and includes an explanation, source metadata, and a direct source URL when available. Data The dataset is provided as JSON Lines: optimot.jsonl Each row contains: Fitxa: Optimot card identifier.… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/optimot-linguistic-data.textquestion-answering1K<n<10K0 likes74 downloads3mo agoHugging Face15softh /alt-kotlin-source-1.4kktext1M<n<10M1 likes72 downloads2y agoHugging Face16xiaobo6668 /math-soft-tokens Math Soft Tokens Dataset Contains training steps: numinamath15_step_11_fixed. text10K<n<100K0 likes59 downloads9mo agoHugging Face17softisight-ai /gbag-bench GBAG-Bench — Grounded BI Answer Generation A public benchmark for the step after the SQL: how faithfully an LLM interprets a query result into a natural-language answer. NL2SQL measures half the problem. GBAG measures the other half. 📂 GitHub (harness, judge, leaderboard): softisight/gbag-bench 📊 Live leaderboard: LEADERBOARD.md 📐 Metric & rubric: METRIC.md 🪪 License: MIT (questions & harness) — bundled SQLite samples retain their original licenses Why this… See the full description on the dataset page: https://huggingface.co/datasets/softisight-ai/gbag-bench.texttable-question-answeringn<1K1 likes49 downloads2mo agoHugging Face18rafidirtiza /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/rafidirtiza/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes44 downloads3mo agoHugging Face19teaql /agentic-software-conformance TeaQL Agentic Software Conformance Machine-readable evidence for the TeaQL Harness: semantic-model evaluation, generated artifacts, seven language-native runtimes, executable examples, and cross-language conformance checks. This is an evidence dataset, not a leaderboard and not a collection of unverified model claims. Each row identifies its evidence level, exact source, verification date, revisions where available, command or gate, result, and important qualifications. The… See the full description on the dataset page: https://huggingface.co/datasets/teaql/agentic-software-conformance.texttext-generationn<1K0 likes44 downloads20d agoHugging Face20robworks-software /texas-k12-curriculum-standards-teks Texas K-12 Curriculum Standards (TEKS-derived) 15,040 generated learning-objective records organized around the Texas Essential Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical Education clusters, and specialized program areas. How this was built (read this first) These records are programmatically generated, not transcribed from official standards documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.texttext-classification10K<n<100K0 likes43 downloads2mo agoHugging Face21tomyimkc /repro-softmax-as-linear-attention-in-the-large-prompt-regime-a-measure-based-perspecti-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes39 downloads2mo agoHugging Face22srcworks-software /nanoset Sourceworks NanoSet NanoSet is an experimental dataset where the main goal is to create a usable chatbot through less training data. What is in NanoSet? NanoSet is divded into 3 major sections, containg 36 entries divided into 6 sub-topics. The structure creates 108 total lines of training data, which may be subject to change in the future. The following is a visual on the structure: 108 entries total 3 Sections, each with 36 entries: Chat Basics (Greetings, Jokes, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/srcworks-software/nanoset.texttext-generationn<1K0 likes38 downloads1y agoHugging Face23referencesource /file-format-software-version-compatibility GIS vector format capabilities and limitations in GDAL Canonical, always-current version: https://referencesource.org/file-format-software-version-compatibility/ Machine-readable: https://referencesource.org/file-format-software-version-compatibility/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-04 Stale after: 2027-01-31 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 3 Key capabilities… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/file-format-software-version-compatibility.textn<1K0 likes35 downloads1mo agoHugging Face24referencesource /software-end-of-support Software end-of-support dates Canonical, always-current version: https://referencesource.org/software-end-of-support/ Machine-readable: https://referencesource.org/software-end-of-support/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-03 Stale after: 2026-11-01 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 68 Release, end-of-active-support and end-of-security-support dates per major version of… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/software-end-of-support.textn<1K0 likes30 downloads1mo agoHugging Face25referencesource /software-interface-version-compatibility Software interface version compatibility Canonical, always-current version: https://referencesource.org/software-interface-version-compatibility/ Machine-readable: https://referencesource.org/software-interface-version-compatibility/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-11 Stale after: 2026-10-10 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 248 Which versions of common software… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/software-interface-version-compatibility.textn<1K0 likes30 downloads1mo agoHugging Face26shovelingpig /soft-label-qnlitabular100K<n<1M0 likes29 downloads2y agoHugging Face27CircularBalls /tt640c-seedsweep-soft-final-s29-v1 TT639G Recombined Tiny Assistant v1 Recombines isolated proof rungs: TT638D code behavior + dyadic/Mercy proof upstream TT639E2 context-copy behavior TT639F3 task-routing behavior simple rule/Q&A behavior Blocking dense gates: seen_combined_pass upstream_regression_pass mixed_heldout_pass anti_collision_pass Do not run dyadic/Mercy compare unless all four gates pass. text100K<n<1M0 likes25 downloads3mo agoHugging Face28softjapan /jaquad-sft softjapan/jaquad-sft データセットの概要 このデータセットは、JaQuAD(Japanese Question Answering Dataset)をSFT(Supervised Fine-Tuning)形式に変換したものです。日本語の質問応答タスクに特化したinstruction tuning用のデータセットです。 データセットの詳細 言語: 日本語 タスク: 質問応答、instruction tuning 形式: SFT(instruction/input/output) 訓練データ: 31,748件 検証データ: 3,939件 合計: 35,687件 データ形式 各サンプルは以下の形式で構成されています: { "id": "tr-000-00-000", "instruction": "次の文脈に基づいて質問に答えてください。可能なら短く正確に答えてください。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/softjapan/jaquad-sft.textquestion-answering10K<n<100K0 likes23 downloads1y agoHugging Face29AgPerry /Video-R1-soft-filter-v2 Video-R1-soft-filter-v2 147,850 samples from Video-R1-260k filtered using the complete Kelsey soft filter. Filter Logic Remove a sample only if 2+ models can answer it text-only under circular evaluation. Model Method MCQ Coverage Non-MCQ GPT-5-mini Single-pass text-only 263K (full) ✅ Qwen2.5-VL-7B 4-perm circular eval 149,586 MCQ pass@10 Gemini 3.1 Pro 3-perm circular eval 149,583 MCQ direct eval (111K) Comparison Version Qwen… See the full description on the dataset page: https://huggingface.co/datasets/AgPerry/Video-R1-soft-filter-v2.textvisual-question-answering100K<n<1M0 likes20 downloads7mo agoHugging Face30monodox /software-engineering-and-devopstextn<1K0 likes20 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.