CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pthinc /BCE-Prettybird-Nano-Themis-v0.1 BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi- Law Dataset (400 Examples) BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi-Law Dataset is a 400-example synthetic dataset developed by Prometech AŞ for experimentation with legal reasoning, instruction following, structured generation, and multi-dimensional response evaluation. Each example combines a task-specific instruction with structured reasoning and quality metadata, including BCE signals, truth and quality values… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Themis-v0.1.texttext-generationn<1K0 likes66 downloads7d agoHugging Face02pthinc /BCE-Prettybird-Nano-OWL-v0.1 BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.texttext-classificationn<1K0 likes59 downloads5mo agoHugging Face03pthinc /BCE-Prettybird-Nano-Hephaistos-v0.1 BCE-Prettybird-Nano-Hephaistos-v0.1 - 1390 Robotics for Instruction-Based Learning BCE-Prettybird-Nano-Hephaistos-v0.1 – 1390 Robotics for Instruction-Based Learning is a bilingual Turkish–English math, sensor, robotics, and embedded-systems QA dataset designed for instruction-based learning, small language models, edge AI research, and robotics education. The dataset focuses on foundational and applied robotics topics such as linear and circular motion, forward and inverse… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Hephaistos-v0.1.texttext-classification1K<n<10K0 likes59 downloads4mo agoHugging Face04pthinc /BCE-Prettybird-Nano-Ulgen-v0.1 BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi- Trader Dataset (320 Examples) BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi-Trader Dataset (320 Examples) is a bilingual Turkish-English synthetic financial reasoning dataset containing 320 instruction-response examples designed for training and evaluating AI systems on investment, portfolio management, corporate finance, risk management, market instruments, valuation, and algorithmic trading tasks. The dataset covers capital… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Ulgen-v0.1.texttext-generationn<1K0 likes49 downloads7d agoHugging Face05pthinc /BCE-Prettybird-Nano-Math-v0.1 BCE-Prettybird-Nano-Math-v0.1 - 500 Math Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math dataset containing 500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability, and… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Math-v0.1.texttext-classificationn<1K0 likes38 downloads6mo agoHugging Face06pthinc /BCE-Prettybird-Nano-Parrot-v0.2 BCE-Prettybird-Nano-Parrot-v0.2 - 700 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.2.texttext-classificationn<1K0 likes33 downloads4mo agoHugging Face07pthinc /BCE-Prettybird-Nano-Science-v0.1 BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.texttext-classificationn<1K0 likes31 downloads6mo agoHugging Face08fs90 /nano-start-data Nano-Start Learning Dataset A small educational dataset for learning how to train language models from scratch. Dataset Description This dataset contains simple, factual examples designed to demonstrate LLM training concepts: Completions: Factual statements the model learns to continue Q&A: Question-answer pairs using chat special tokens Chat: Multi-turn conversations with system prompts The dataset is intentionally small (~276 examples) so models can be trained quickly… See the full description on the dataset page: https://huggingface.co/datasets/fs90/nano-start-data.texttext-generationn<1K0 likes30 downloads10mo agoHugging Face09pthinc /BCE-Prettybird-Nano-Apollo-v0.1 BCE-Prettybird-Nano-Apollo-v0.1 Synthetic Multi-Language Software Engineering & UI/UX Dataset (1,070 Examples) This dataset contains 1,070 synthetic, high-quality examples covering a broad range of software engineering, architecture, database development, web design, UI/UX design, and design pattern implementations across multiple programming languages and frameworks. The collection includes: SOLID principle code examples in PHP, C#, Python, C++, Java, and JavaScript Design… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apollo-v0.1.texttext-generation1K<n<10K0 likes29 downloads4mo agoHugging Face10pthinc /BCE-Prettybird-Nano-Kangal-v0.1 BCE-Prettybird-Nano-Kangal-v0.1 - 525 LOVE Q&A Dataset for Instruction-Based Learning The "BCE-Prettybird-Nano-Kangal-v0.1: Love Dataset" consists of 525 rows of insightful data, offering a comprehensive exploration of romantic relationships. Covering diverse aspects from sexuality and intimacy to romance, family life management, and tips on how to treat women, this dataset delves into the complexities of modern relationships. It aims to provide valuable perspectives for those… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kangal-v0.1.texttext-classificationn<1K0 likes28 downloads24d agoHugging Face11pthinc /BCE-Prettybird-Nano-Kayra-v0.1 BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.tabulartext-classificationn<1K0 likes28 downloads4mo agoHugging Face12pthinc /BCE-Prettybird-Nano-Merkur-v0.1 BCE-Prettybird-Nano-Merkur-v0.1 - 4300 Chatting for Instruction-Based Learning BCE-Prettybird-Nano-Merkur-v0.1 – This dataset contains 4,300 bilingual Turkish-English conversational dialogue samples designed for training and fine-tuning conversational AI systems, chatbots, and large language models. The dataset includes four carefully curated categories: Flirty General Chat, featuring playful and socially engaging conversations; Polite General Conversation, focused on respectful… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Merkur-v0.1.texttext-classification1K<n<10K0 likes25 downloads4mo agoHugging Face13pthinc /BCE-Prettybird-Nano-Thoth-v0.1 BCE-Prettybird-Nano-Thoth-v0.1 320 Latex Katex Math Q&A Dataset for Instruction-Based Learning BCE-Prettybird-Nano-Thoth-v0.1 is a compact KaTeX/LaTeX-oriented BCE dataset released under pthinc/BCE-Prettybird-Nano-Thoth-v0.1, designed to explore how small behavioral-control datasets can teach structured mathematical expression, symbolic formatting, and explanation consistency in LLM workflows. The dataset contains 320 question–answer pairs focused on KaTeX and LaTeX usage, covering… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Thoth-v0.1.texttext-classificationn<1K0 likes23 downloads4mo agoHugging Face14castorini /NanoKnow_Benchmark NanoKnow Benchmark [Paper] [Code] Pre-built NanoKnow projections of SQuAD and Natural Questions (NQ-Open) onto karpathy/fineweb-edu-100b-shuffle, the pre-training corpus used by nanochat. NanoKnow partitions benchmark questions into: Supported: at least one answer-bearing pre-training document was verified by an LLM judge. Unsupported: no answer-bearing pre-training document was verified. Data summary Dataset Questions Supported Unsupported Supported qrels… See the full description on the dataset page: https://huggingface.co/datasets/castorini/NanoKnow_Benchmark.textquestion-answering10K<n<100K0 likes20 downloads2mo agoHugging Face15pthinc /BCE-Prettybird-Nano-Parrot-v0.1 BCE-Prettybird-Nano-Parrot-v0.1 - 200 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.1.texttext-classificationn<1K1 likes20 downloads4mo agoHugging Face16wflying /nemotron-nano-rl-mcqa-19k Nemotron Nano RL MCQA 19K Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format. Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.textquestion-answering10K<n<100K0 likes20 downloads2mo agoHugging Face17pthinc /BCE-Prettybird-Nano-Apep-v0.1 BCE-Prettybird-Nano-Apep-v0.1 - 1025 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apep-v0.1.texttext-classification1K<n<10K0 likes16 downloads5mo agoHugging Face18alxraun /ers-bench-nano ERS-Bench-Nano A parallel QA dataset designed to evaluate the effects of ERS semantic compression by comparing answer accuracy between natural language source text and its lossy LLM-generated ERS artifact. Schema Field Type Description domain string Category: code, formal, science, humanities. sample_id string Sample identifier. source string Original natural language text. ers_artifact string ERS notation artifact. question_id string Query identifier.… See the full description on the dataset page: https://huggingface.co/datasets/alxraun/ers-bench-nano.textquestion-answeringn<1K0 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.