CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jtregunna /software-strategist-v1 Software Fundamentals — Strategy Knowledge Base A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists. The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON. Dataset Summary This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.texttext-generation10K<n<100K2 likes155 downloads4mo agoHugging Face02robworks-software /k12-standards-instruction-tasks K-12 Curriculum Tasks (generated) 2,489 generated instruction/input/output records covering five curriculum tasks: assessment creation, learning objective generation, misconception detection, standard explanation, and standards Q&A. Content is predominantly mathematics. Important: the name is misleading Despite the name, this dataset contains no school directory data. There are four columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.texttext-generation1K<n<10K1 likes146 downloads2mo agoHugging Face03softcatala /mantinc-catalan-drift Mantinc — Catalan Drift Benchmark Descripció (ca) Mantinc és un banc de proves que avalua si un model de llenguatge continua responent en català quan el missatge, la conversa prèvia o el context recuperat l'empenyen a fer-ho en una altra llengua, normalment el castellà o l'anglès. Dataset Description Mantinc is a benchmark that measures whether a language model keeps answering in Catalan when the prompt, prior conversation, or retrieved context… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/mantinc-catalan-drift.texttext-generationn<1K0 likes128 downloads18d agoHugging Face04softisight-ai /gbag-bench GBAG-Bench — Grounded BI Answer Generation A public benchmark for the step after the SQL: how faithfully an LLM interprets a query result into a natural-language answer. NL2SQL measures half the problem. GBAG measures the other half. 📂 GitHub (harness, judge, leaderboard): softisight/gbag-bench 📊 Live leaderboard: LEADERBOARD.md 📐 Metric & rubric: METRIC.md 🪪 License: MIT (questions & harness) — bundled SQLite samples retain their original licenses Why this… See the full description on the dataset page: https://huggingface.co/datasets/softisight-ai/gbag-bench.texttable-question-answeringn<1K1 likes49 downloads2mo agoHugging Face05teaql /agentic-software-conformance TeaQL Agentic Software Conformance Machine-readable evidence for the TeaQL Harness: semantic-model evaluation, generated artifacts, seven language-native runtimes, executable examples, and cross-language conformance checks. This is an evidence dataset, not a leaderboard and not a collection of unverified model claims. Each row identifies its evidence level, exact source, verification date, revisions where available, command or gate, result, and important qualifications. The… See the full description on the dataset page: https://huggingface.co/datasets/teaql/agentic-software-conformance.texttext-generationn<1K0 likes44 downloads20d agoHugging Face06robworks-software /texas-k12-curriculum-standards-teks Texas K-12 Curriculum Standards (TEKS-derived) 15,040 generated learning-objective records organized around the Texas Essential Knowledge and Skills (TEKS) taxonomy, spanning core academic subjects, Career & Technical Education clusters, and specialized program areas. How this was built (read this first) These records are programmatically generated, not transcribed from official standards documents. A generator took a standards taxonomy - codes, grade levels… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/texas-k12-curriculum-standards-teks.texttext-classification10K<n<100K0 likes43 downloads2mo agoHugging Face07srcworks-software /nanoset Sourceworks NanoSet NanoSet is an experimental dataset where the main goal is to create a usable chatbot through less training data. What is in NanoSet? NanoSet is divded into 3 major sections, containg 36 entries divided into 6 sub-topics. The structure creates 108 total lines of training data, which may be subject to change in the future. The following is a visual on the structure: 108 entries total 3 Sections, each with 36 entries: Chat Basics (Greetings, Jokes, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/srcworks-software/nanoset.texttext-generationn<1K0 likes38 downloads1y agoHugging Face08softjapan /jaquad-sft softjapan/jaquad-sft データセットの概要 このデータセットは、JaQuAD(Japanese Question Answering Dataset)をSFT(Supervised Fine-Tuning)形式に変換したものです。日本語の質問応答タスクに特化したinstruction tuning用のデータセットです。 データセットの詳細 言語: 日本語 タスク: 質問応答、instruction tuning 形式: SFT(instruction/input/output) 訓練データ: 31,748件 検証データ: 3,939件 合計: 35,687件 データ形式 各サンプルは以下の形式で構成されています: { "id": "tr-000-00-000", "instruction": "次の文脈に基づいて質問に答えてください。可能なら短く正確に答えてください。", "input":… See the full description on the dataset page: https://huggingface.co/datasets/softjapan/jaquad-sft.textquestion-answering10K<n<100K0 likes23 downloads1y agoHugging Face09Inkwell-Software /screenplay-revision-evaluation Screenplay Revision Evaluation Cases 24 original screenwriting revision tasks. Each gives a short scene and a constraint — cut a page to its beat, plant a prop, hold an answer back, fix a continuity slip — then pairs it with mechanical checks (a word ceiling, a line that must survive) and separate human-review questions. It tests whether a tool, or a person, can make a tightly-constrained edit while keeping the scene intact. Each task's reference_output is null, because a… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-revision-evaluation.texttext-generationn<1K0 likes7h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.