CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kingkw1 /read-along-ai-agent-traces Read-Along AI - Agent Traces This dataset contains the raw agent traces and conversation logs from the development of Read-Along AI, a submission for the Hugging Face Build Small Hackathon. Dataset Description These .jsonl files represent the unedited, behind-the-scenes "agent traces" of the AI coding assistant orchestrating the build of this project. Sharing these traces fulfills the requirements for the "Sharing is Caring" bonus badge, providing the community… See the full description on the dataset page: https://huggingface.co/datasets/kingkw1/read-along-ai-agent-traces.tabulartext-generationn<1K0 likes176 downloads3mo agoHugging Face02mbazaNLP /kinyarwanda_monolingual_v01.0 !!! PLEASE USE mbazaNLP/kinyarwanda_monolingual_v01.1 !!! !!! This version contains several duplicates and few non-kinyarwanda documents Dataset Summary The Kinyarwanda Monolingual Dataset version 1 is a large collection of Kinyarwanda language texts aimed at supporting the development of NLP and AI applications which can process Kinyarwanda texts. This dataset contains 78k documents, totalling about 25 million words, and includes diverse content types such… See the full description on the dataset page: https://huggingface.co/datasets/mbazaNLP/kinyarwanda_monolingual_v01.0.tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face03kingkaung /english_islamqainfo Dataset Card for English Islam QA Info Dataset Description The English Islam QA Info (19,052 questions and answers) is derived from the IslamQA website and contains curated question-and-answer pairs categorized by topic. It serves as a resource for multilingual and cross-lingual natural language processing (NLP) tasks. This dataset is part of a broader initiative to enhance the understanding and computational handling of Islamic jurisprudence and advice. Key… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/english_islamqainfo.tabulartable-question-answering10K<n<100K5 likes24 downloads2y agoHugging Face04kingkaung /Quran_English_Myanmar_Parrelel_Corpus Quran English-Myanmar Parallel Corpus Description This dataset is a parallel corpus of the Quran, containing translations in English and Myanmar. It includes 6,237 verses (ayahs) from all chapters (surahs), aligned by their respective Surah and Ayah numbers. English Translation: Provided by Dr. Muhsin Khan and Dr. Hilali. Myanmar Translation: Translated by the Myanmar Quran Translation Committee, comprising religious and non-religious scholars, and later published by… See the full description on the dataset page: https://huggingface.co/datasets/kingkaung/Quran_English_Myanmar_Parrelel_Corpus.tabulartranslation1K<n<10K0 likes22 downloads2y agoHugging Face05build-small-hackathon /Kintsugi-Garden-traces Kintsugi Garden Evaluation Traces Paired evaluation traces from Kintsugi Garden — a local-first Jungian dream journal that runs Qwen3-8B through llama.cpp on a ZeroGPU Space. Every entry the app produces is shaped by both a fine-tuned model and a four-layer voice/safety architecture; this dataset is what those layers look like under instrumentation. What's in here 114 deterministic runs over the same 19 prompts × 3 trials, evenly split between: baseline —… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/Kintsugi-Garden-traces.tabulartext-generationn<1K0 likes13 downloads4mo agoHugging Face06kinit /synthetic-queries-and-ml-instructionsgated Synthetic Dataset: Queries and ML Instructions Dataset Description Dataset Summary This is a synthetic dataset with queries in Slovak language and ML instructions. The dataset was designed to train a model for extracting structured machine learning task requirements from natural language user queries. The dataset contains user queries in Slovak describing ML tasks paired with structured JSON outputs containing task attributes like dataset modality, task type… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-queries-and-ml-instructions.tabulartext-generation10K<n<100K1 likes6 downloads11mo agoHugging Face07kinit /synthetic-conversations-and-ml-instructionsgated Synthetic Dataset: Conversations and ML Instructions Dataset Description Dataset Summary This is a synthetic dataset of 5,000 Slovak multi-turn ML advisory conversations paired with structured JSON outputs. The dataset was designed to train and evaluate models that extract machine learning task requirements from realistic, natural Slovak dialogue, including cases where requirements are revealed gradually or changed during the conversation. Each… See the full description on the dataset page: https://huggingface.co/datasets/kinit/synthetic-conversations-and-ml-instructions.tabulartext-generation1K<n<10K1 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.