CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bbidpa /flutter-diff-steps-v1 Flutter Codegen: Diff Steps Synthetic dataset of step-by-step Flutter/Dart widget construction, where each row is one incremental edit in a sequence: given a goal, the current code, and the history of steps taken so far, predict the next action (a short description) and the code change as a search/replace diff hunk. Built for training and evaluating small language models on iterative, diff-based code editing -- as opposed to regenerating the whole file at each step. This is the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-diff-steps-v1.tabulartext-generation100K<n<1M0 likes177 downloads17d agoHugging Face02bbidpa /flutter-full-examples-v1 Flutter Codegen: Full Examples Synthetic dataset of complete Flutter/Dart widgets, each paired with the goal that describes them and (optionally) starting code. Unlike flutter-codegen-diff-steps, there's no step history or diff structure here -- each row is a single, standalone goal -> complete file example. This is the whole-code counterpart to flutter-diff-steps-v1, intended for training/evaluating a baseline that generates the entire file in one shot, to compare against the… See the full description on the dataset page: https://huggingface.co/datasets/bbidpa/flutter-full-examples-v1.tabulartext-generation10K<n<100K0 likes136 downloads17d agoHugging Face03flufy3d /xinhe-dataset Xinhe Memory Training Dataset 中文长上下文 + 多轮对话训练数据,用于 Xinhe (心核) 项目让小型 Transformer 在 fast-weights NeuralMemory 中习得长程记忆能力 (v9.5 paper-faithful Titans MAC + LoRA + per-layer K/V)。 包含五个 config: skeleton:11 种合成骨架对话(needle-in-haystack 写读 / 覆写 / 删除 / stale-read 对抗), v9.5 起 S1/S2/S5/S7/S9/S10/S11 切到 paragraph distract(从 congliu 短答 bank 拼成长段), 用于 sanity probe + 长 episode 压力 dialog:LLM(DeepSeek / OpenRouter)生成的 5-Beat 自然多轮对话,真实分布代理 novel:中文长篇小说按 ~350 字 raw chunking,每 episode = 8 个章内连续 chunk… See the full description on the dataset page: https://huggingface.co/datasets/flufy3d/xinhe-dataset.texttext-generation100K<n<1M2 likes44 downloads5mo agoHugging Face04KavinduHansaka /prompt-gen-10k-flux-sdxl Prompt Generation Dataset (10K Narrative for Flux / SDXL) This dataset (prompt_gen_final_10k.jsonl and prompt_gen_final_10k.csv) was used to train and fine-tune image-prompt models such as KavinduHansaka/Llama-3.2-1B-ImageGen. It contains 10,000 curated narrative prompt samples designed for image generation models like Stable Diffusion XL and Flux.Unlike raw tag-based datasets, the target field provides natural paragraphs (≈80–100 words) that describe cinematic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/KavinduHansaka/prompt-gen-10k-flux-sdxl.tabulartext-generation10K<n<100K0 likes43 downloads1y agoHugging Face05NoirZangetsu /flutterft FlutterFT FlutterFT is a 5,000-record English-language instruction-tuning corpus for Flutter/Dart code generation, spanning seven task types: question-to-code, test generation, bug fixing, refactoring, completion, code explanation, and API usage. Each record is a chat-format (system/user/assistant) example. Important notice on source provenance and licensing Source repository, commit, and license information for the underlying code snippets was not recorded during… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/flutterft.texttext-generation1K<n<10K0 likes43 downloads2mo agoHugging Face06arshiaafshani /persian-natural-fluently Persian scientific dataset I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology. The content of the dataset has been generated by : Human, Grok3, DeepSeek R1. License This dataset is licensed under apache-2.0. texttext-generationn<1K14 likes39 downloads1y agoHugging Face07zshiyi /Fluid-Mechanics-CoT 🌊 Engineering Fluid Mechanics CoT Dataset (工程流体力学思维链数据集) 📖 Dataset Description (数据集简介) This dataset focuses on Engineering Fluid Mechanics, specifically designed to enhance Large Language Models' (LLMs) reasoning capabilities in complex physics problems. Unlike standard QA datasets, this dataset provides Chain-of-Thought (CoT) annotations, breaking down the problem-solving process into: Analysis & Reasoning: Strategy selection and physical law identification.… See the full description on the dataset page: https://huggingface.co/datasets/zshiyi/Fluid-Mechanics-CoT.textquestion-answeringn<1K1 likes19 downloads9mo agoHugging Face08newyouth19 /jfleg_fluencygatedtexttext-generation1K<n<10K0 likes1 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.