CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01b-mc2 /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.texttext-generation10K<n<100K506 likes6.3k downloads3y agoHugging Face02yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.2k downloads7mo agoHugging Face03rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes511 downloads6mo agoHugging Face04kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes297 downloads6mo agoHugging Face05windchimeran /creativemath_fulltabularquestion-answering1K<n<10K0 likes123 downloads1y agoHugging Face06creeperdatasets /python_debugging Python Debugging A synthetic instruction-tuning dataset for training AI models to identify and fix bugs in Python code. Dataset Summary Field Value Entries 75 Format input / output pairs Language English Topic Finding and fixing bugs in Python code Synthetic Yes, generated with DeepSeek License MIT Dataset Description Each entry presents a snippet of Python code containing a deliberate bug, along with a corrected version… See the full description on the dataset page: https://huggingface.co/datasets/creeperdatasets/python_debugging.texttext-generationn<1K0 likes92 downloads29d agoHugging Face07DarkyMan /Opus-4.6-RU-Reasoning-creative-1385x-not-filtered Opus-4.6-RU-Creative-Writing — Russian Creative Writing Reasoning Dataset A Russian-language dataset of creative writing tasks generated with Claude claude-opus-4.6 (extended thinking enabled). Each sample contains a creative prompt, a full reasoning chain showing the creative process, and a detailed artistic response. Dataset Info Language: Russian 🇷🇺 Size: ~1,385 samples (growing) Model used: anthropic/claude-opus-4.6 with reasoning: {effort: "high"} Format:… See the full description on the dataset page: https://huggingface.co/datasets/DarkyMan/Opus-4.6-RU-Reasoning-creative-1385x-not-filtered.texttext-generation1K<n<10K3 likes88 downloads6mo agoHugging Face08mismayil /cresowlve CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge Dataset Description This is a bilingual benchmark for creative problem-solving grounded in real-world knowledge and solvable by human experts. CresOWLve spans a diverse range of knowledge and creative domains, varies in difficulty, requires multiple creative thinking strategies, and is manually validated to ensure quality. It contains ~2K open-ended questions with answers and explanations.… See the full description on the dataset page: https://huggingface.co/datasets/mismayil/cresowlve.textquestion-answering1K<n<10K1 likes73 downloads3mo agoHugging Face09emdemor /sql-create-context-pt Overview Este dataset é uma versão traduzida para o português do dataset b-mc2/sql-create-context, que foi construído a partir dos datasets WikiSQL e Spider. Ele contém exemplos de perguntas em português, instruções SQL CREATE TABLE e consultas SQL que respondem às perguntas utilizando a instrução CREATE TABLE como contexto. O principal objetivo deste dataset é ajudar modelos de linguagem natural em português a gerar consultas SQL precisas e contextualizadas, prevenindo a… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/sql-create-context-pt.texttext-generation10K<n<100K2 likes71 downloads2y agoHugging Face10bugdaryan /sql-create-context-instruction Overview This dataset is built upon SQL Create Context, which in turn was constructed using data from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-SQL LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-SQL datasets. The CREATE TABLE statement can often be… See the full description on the dataset page: https://huggingface.co/datasets/bugdaryan/sql-create-context-instruction.texttext-generation10K<n<100K19 likes68 downloads3y agoHugging Face11philschmid /sql-create-context-copy Fork of b-mc2/sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.texttext-generation10K<n<100K4 likes65 downloads3y agoHugging Face12detakarang /sql-create-context-id Overview This dataset is a fork from sql-create-context This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/detakarang/sql-create-context-id.texttext-generation10K<n<100K0 likes57 downloads3y agoHugging Face13hiltch /pandas-create-context Overview This dataset is built from sql-create-context, which in itself builds from WikiSQL and Spider. I have used GPT4 to translate the SQL schema into pandas DataFrame schem initialization statements and to translate the SQL queries into pandas queries. There are 862 examples of natural language queries, pandas DataFrame creation statements, and pandas query answering the question using the DataFrame creation statement as context. This dataset was built with text-to-pandas… See the full description on the dataset page: https://huggingface.co/datasets/hiltch/pandas-create-context.texttext-generation10K<n<100K2 likes57 downloads3y agoHugging Face14msw-ai-tf /maplestory-worlds-creator-qa MapleStory Worlds Creator QA Synthetic question-answer dataset built from the official MapleStory Worlds Creator Center documentation. Questions are generated to be self-contained and grounded in the source docs; answers avoid source/meta references so they read like an expert explanation. Some QA pairs are composed from multiple related documents (see combo_sources). Parallel Korean/English. Intended for instruction tuning, QA, and retrieval. Composition… See the full description on the dataset page: https://huggingface.co/datasets/msw-ai-tf/maplestory-worlds-creator-qa.textquestion-answering100K<n<1M2 likes56 downloads3mo agoHugging Face15marcuscedricridia /Qwill-RP-CreativeWriting-Reasoning Qwill RP CreativeWriting Reasoning Dataset 📝 Dataset Summary Qwill-RP-CreativeWriting-Reasoning is a creative writing dataset focused on structured reasoning. Each row contains a fictional or narrative prompt sourced from nothingiisreal/Reddit-Dirty-And-WritingPrompts, along with an AI-generated response that includes: Reasoning, wrapped in <think>...</think> Final Answer, wrapped in <answer>...</answer> The goal is to train or evaluate models on chain-of-thought… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/Qwill-RP-CreativeWriting-Reasoning.tabulartext-generation1K<n<10K8 likes53 downloads1y agoHugging Face16MaltbyTom /CREATE-Protocol CREATE Protocol: Cognitive Recursion Enhancement for Applied Transform Evolution Dataset Description CREATE (Cognitive Recursion Enhancement for Applied Transform Evolution) is a structured cognitive scaffolding framework designed to support epistemic integrity, curiosity-driven inquiry, and aligned reasoning in both human and artificial cognitive systems. The protocol consists of modular text packets that provide frameworks for navigating uncertainty, recognizing… See the full description on the dataset page: https://huggingface.co/datasets/MaltbyTom/CREATE-Protocol.texttext-generationn<1K0 likes42 downloads8mo agoHugging Face17michsethowusu /Code-170k-mauritian-creole Dataset Description Code-170k-mauritian-creole is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Mauritian Creole, making coding education accessible to Mauritian Creole speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Mauritian Creole language - democratizing coding education Multi-turn dialogues covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-mauritian-creole.texttext-generation100K<n<1M0 likes41 downloads11mo agoHugging Face18CreativeAlloyYT /French_Grammar_Explanations This dataset contains 1500+ French grammar explanations. It's the one I used to train my finetuned LLM called FrenchLlama-3.2-1B-Instruct. You can use this dataset for your own training purposes & find the aforementioned model on my HuggingFace profile. textquestion-answering1K<n<10K0 likes36 downloads1y agoHugging Face19crevious /Tri-NL Tri-NL: A Multi-Domain Reasoning Dataset A synthetic reasoning dataset spanning Mathematics, Coding, and Healthcare — designed for training and evaluating LLMs on step-by-step reasoning across domains. Key Features 1,500 samples — 500 per domain (math, coding, healthcare) Step-by-step reasoning — every answer follows a numbered Step 1: → Final Answer: format Cross-domain overlap — ~15% of each domain intentionally overlaps with each other domain (e.g., biostatistics… See the full description on the dataset page: https://huggingface.co/datasets/crevious/Tri-NL.textquestion-answering1K<n<10K1 likes31 downloads7mo agoHugging Face20saksornr /sql-create-context-thai Overview This dataset builds from sql-create-context. @misc{b-mc2_2023_sql-create-context, title = {sql-create-context Dataset}, author = {b-mc2}, year = {2023}, url = {https://huggingface.co/datasets/b-mc2/sql-create-context}, note = {This dataset was created by modifying data from the following sources: \cite{zhongSeq2SQL2017, yu2018spider}.}, } texttext-generation10K<n<100K0 likes30 downloads2y agoHugging Face21gtfintechlab /CreditQA CCA Numerical Reasoning CCA Numerical Reasoning is an 800-question dataset of credit-card-agreement questions designed to evaluate numerical reasoning over consumer finance terms. Each example contains a natural-language question, a ground-truth answer, output units, input-unit metadata, executable-style reasoning steps, a program trace, question tense metadata, financial-term tags, and the source credit card agreement identifier. The dataset contains 800 numerical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/gtfintechlab/CreditQA.textquestion-answeringn<1K0 likes29 downloads2mo agoHugging Face22DevchandraSah /indian-credit-card-facts Indian Credit Card Facts — an open dataset A dated, machine-readable record of the Indian credit-card market: what each card charges, what it earns, how its terms have changed over time, where its points can be transferred, and what those points are worth. Four tables, JSON and CSV, CC BY 4.0. Every release is immutable and separately citable. Table Rows Unit As of cards 280 cards 2026-08-15 changes 1617 events 2026-08-13 transfers 153 edges 2026-07-02… See the full description on the dataset page: https://huggingface.co/datasets/DevchandraSah/indian-credit-card-facts.texttabular-classification1K<n<10K0 likes24 downloads1mo agoHugging Face23michsethowusu /Code-170k-seychellois-creole Dataset Description Code-170k-seychellois-creole is a groundbreaking dataset containing 176,999 programming conversations, originally sourced from glaiveai/glaive-code-assistant-v2 and translated into Seychellois Creole, making coding education accessible to Seychellois Creole speakers. 🌟 Key Features 176,999 high-quality conversations about programming and coding Pure Seychellois Creole language - democratizing coding education Multi-turn dialogues covering various… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/Code-170k-seychellois-creole.texttext-generation100K<n<1M0 likes20 downloads11mo agoHugging Face24dipanjanS /sql-create-context Overview This dataset builds from WikiSQL and Spider. There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/dipanjanS/sql-create-context.texttext-generation10K<n<100K0 likes18 downloads6mo agoHugging Face25yuneun92 /create_qa_news 질문 생성: kullm3 모델 이용 답변 생성: GPT3.5 turbo API 이용 지문 원본: AI HUB 뉴스 기계독해 데이터셋 textquestion-answering1K<n<10K0 likes14 downloads2y agoHugging Face26creeperdatasets /qemu_networking Qemu Networking from Claude Haiku 4.5 A synthetic instruction-tuning dataset covering QEMU networking concepts, generated using Claude Haiku 4.5. Dataset Summary Total rows: 75 Topic: QEMU virtual networking Difficulty distribution: Easy, Intermediate, Advanced 278 unique tags across networking subtopics Splits train: 75 rows Columns id: Stable entry ID instruction: Instruction text for fine-tuning input: Original prompt/question output:… See the full description on the dataset page: https://huggingface.co/datasets/creeperdatasets/qemu_networking.texttext-generationn<1K0 likes11 downloads5mo agoHugging Face27chengq9 /CreativityBench-MM MM-CreativityBench MM-CreativityBench evaluates creative tool-repurposing in multimodal models through visual, part-level grounding. Splits train: 868 examples test: 333 examples The loadable split files are: data/train.jsonl data/test.jsonl Nested fields from the source JSON files are serialized as JSON strings in columns such as setting_json, golds_json, entities_json, items_json, and solution_json. The original unmodified files are also included under raw/.… See the full description on the dataset page: https://huggingface.co/datasets/chengq9/CreativityBench-MM.textquestion-answering1K<n<10K0 likes9 downloads4mo agoHugging Face28cfosilvia /silvia-creditcard-bench Silvia Credit Card Bench (v1.0) Silvia Credit Card Bench is a benchmark of 100 personalized credit-card advisory scenarios. Each scenario pairs a realistic, anonymized user financial profile (including their existing credit-card portfolio) with a question about card selection, rewards optimization, and spend routing. Every scenario carries per-configuration judge scores for answer accuracy and grounding. Schema One JSONL file (data/train.jsonl), config default… See the full description on the dataset page: https://huggingface.co/datasets/cfosilvia/silvia-creditcard-bench.textquestion-answeringn<1K0 likes7 downloads10d agoHugging Face29person65 /Gemma3_4b-Created-Chats This dataset is AI generated, the creator might have missed errors This dataset contains 5110 lines of jsonl. It follows this format: {"messages": [{"role" : "User", "content" : "..."}, {"role" : "Assistant", "content" : "..."}] The dataset contains a variety of topics and a random number of turns per chat, with 14 turns as the maximum. It was generated with Gemma3:4b using ollama with a temperature of 0.8 on a RTX 3050 8GB.in order to download the dataset, use the following code:… See the full description on the dataset page: https://huggingface.co/datasets/person65/Gemma3_4b-Created-Chats.question-answering10K<n<100K0 likes5 downloads8mo agoHugging Face30Ricardo6935 /Creative_LLMquestion-answering10M<n<100M0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.