CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MK4-Research /LOREA-cyber-training-data LOREA-cyber security code-analysis training set Two corpora live here. The v6_corpus config is the newer one and is what actually trained LOREA-cyber v6 Pilot. The eight older configs are the v5-era set, kept as-is because they are a different schema and still useful on their own. v6_corpus 4,780 train and 151 validation rows in chat format: {"messages": [...], "meta": {...}}, where messages is a system/user/assistant sequence and meta carries type, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-training-data.texttext-generation1K<n<10K0 likes99 downloads29d agoHugging Face02AgentZeroCopeAI /lore-corpus COPEAI Lore Corpus Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus, mood label definitions, and disclosure copy from COPEAI — an AI-themed Solana memecoin satire on Pump.fun. Compliance frame: Every entry here is fictional in-character satire. Nothing in this corpus is financial advice, investment guidance, or a recommendation to transact. COPEAI provides no rights, utility, yield, or appreciation expectations. The agent names (TRON, CLU, QUORRA, ZUSE… See the full description on the dataset page: https://huggingface.co/datasets/AgentZeroCopeAI/lore-corpus.texttext-generationn<1K2 likes60 downloads5mo agoHugging Face03lorenzo217 /HundredCV-Chat 百人对话数据集 HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs 简介 本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。 数据集具有如下特点: 自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。 多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。 高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。 数据样例 HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/lorenzo217/HundredCV-Chat.texttext-generation10K<n<100K1 likes51 downloads6mo agoHugging Face04Lo-Renz-O /grade-school-math-instructions-Malagasy Overview This dataset is a Malagasy adaptation of grade-school-math-instructions. It consists of arithmetic word problems converted into instruction-answer pairs in Malagasy. Each entry contains a math problem presented as an instruction, optional contextual input, and a detailed step-by-step solution in Malagasy. The dataset is particularly useful for training and evaluating models on arithmetic reasoning and instruction-following tasks in Malagasy, a low-resource language.… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/grade-school-math-instructions-Malagasy.textquestion-answering1K<n<10K0 likes38 downloads10mo agoHugging Face05Dula23 /lore-corpus COPEAI Lore Corpus Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus, mood label definitions, and disclosure copy from COPEAI — an AI-themed Solana memecoin satire on Pump.fun. Compliance frame: Every entry here is fictional in-character satire. Nothing in this corpus is financial advice, investment guidance, or a recommendation to transact. COPEAI provides no rights, utility, yield, or appreciation expectations. The agent names (TRON, CLU, QUORRA… See the full description on the dataset page: https://huggingface.co/datasets/Dula23/lore-corpus.texttext-generationn<1K0 likes38 downloads7d agoHugging Face06Lo-Renz-O /malagasy-sentence Overview This dataset consists of clean, structured sentences extracted via Optical Character Recognition (OCR) from approximately 1GB of Malagasy thesis documents. These documents were collected based on educational, cultural, and linguistic themes. The dataset is saved in CSV format, and is particularly useful for NLP tasks involving sentence-level modeling in Malagasy — a low-resource language. Dataset Details Language: Malagasy Source: OCR'd academic thesis… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/malagasy-sentence.texttext-generation100K<n<1M0 likes26 downloads1y agoHugging Face07sfrontull /la-usc_valbadia_loresmt24 Dataset Card: Ladin (Val Badia) - Monolingual (La Usc di Ladins) Overview Source Paper: "Rule-Based, Neural and LLM Back-Translation: Comparative Insights from a Variant of Ladin" Description: This dataset contains monolingual sentences in Ladin (Val Badia) from the newspaper "La Usc di Ladins" (https://www.lausc.it), which has been archived since 2008. The newspaper offers texts in five variants of Ladin, corresponding to the five Ladin valleys. We extracted 1,937,608… See the full description on the dataset page: https://huggingface.co/datasets/sfrontull/la-usc_valbadia_loresmt24.texttext-classification100K<n<1M1 likes25 downloads7mo agoHugging Face08Lo-Renz-O /Collection-Tononkalo-Malagasy Collection Tononkalo Malagasy Dataset Description This dataset is a collection of Malagasy poems (Tononkalo) scraped from Vetso Serasera. It focuses on creative writing, rhymes, and artistic expression in the Malagasy language. Important Note: This dataset has been rigorously filtered. Language Filtering: Poems written primarily in French or English have been removed to ensure a high-quality Malagasy corpus. Cleaning: Metadata, dates, author signatures inside the text… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/Collection-Tononkalo-Malagasy.texttext-generation10K<n<100K0 likes19 downloads8mo agoHugging Face09LorenaYannnnn /how_lms_answer_one_to_many_factual_queries One-to-Many Factual Queries Datasets This is the official dataset used in our EMNLP 2025 paper Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries. The dataset includes six subsets named {dataset_name}_template_{i}, where dataset_name is country_cities, artist_songs, or actor_movies, and each dataset has three prompt templates (i = 1, 2, 3). The {model_name}_step_{i} split in each subset contains the data used for analyzing model_name's behavior at… See the full description on the dataset page: https://huggingface.co/datasets/LorenaYannnnn/how_lms_answer_one_to_many_factual_queries.question-answering0 likes17 downloads1y agoHugging Face10TSEOsiris /bannerlord-lore-dataset Bannerlord Lore Dataset Comprehensive lore dataset for Mount & Blade II: Bannerlord - a medieval action RPG by TaleWorlds Entertainment. Dataset Description This dataset contains structured lore information extracted from the game, including: Categories Category Description Languages Heroes NPCs, lords, companions EN, RU, TR Kingdoms Major factions (Empire, Battania, etc.) EN, RU, TR Settlements Cities, castles, villages EN, RU, TR Cultures… See the full description on the dataset page: https://huggingface.co/datasets/TSEOsiris/bannerlord-lore-dataset.texttext-generation1K<n<10K0 likes15 downloads9mo agoHugging Face11ogmatrixllm /pokemon-lore-instructionstexttext-generationn<1K1 likes14 downloads2y agoHugging Face12Lo-Renz-O /Multilingual-Thinking-with-Malagasy Multilingual Thinking with Malagasy (mg) Dataset Summary This dataset is a modified version of HuggingFaceH4/Multilingual-Thinking. In this version, the Spanish subset has been replaced with Malagasy (mg) translations to support Chain-of-Thought (CoT) reasoning for low-resource languages. The dataset retains all other original languages (English, French, German, etc.). The reasoning traces (analysis field) were translated using Google Gemini 2.0 Flash/1.5 Flash, with… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/Multilingual-Thinking-with-Malagasy.texttext-generation1K<n<10K0 likes11 downloads9mo agoHugging Face13Lo-Renz-O /alpaca-gpt4-Malagasygated Overview This dataset is a Malagasy adaptation of the Alpaca-GPT4 instruction-following dataset.It contains instruction-response pairs translated or adapted into Malagasy, designed for fine-tuning instruction-following language models. Each entry includes an instruction, optional input context, and a reference response generated by GPT-4 and adapted to Malagasy using Gemini 2.5 for the translation. The dataset enables training and evaluating LLMs on instruction understanding… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/alpaca-gpt4-Malagasy.textquestion-answering10K<n<100K0 likes9 downloads10mo agoHugging Face14Lo-Renz-O /GBV-Malagasy Dataset Description This dataset contains news articles, stories, and cultural reports scraped from Global Voices Malagasy. It is intended to support Natural Language Processing (NLP) tasks for the Malagasy language, such as language modeling and text generation. The data is automatically scraped and updated once a month to ensure freshness. Dataset Structure The dataset is formatted in JSONL (JSON Lines). Each entry represents a single article. Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/GBV-Malagasy.texttext-generation10K<n<100K0 likes8 downloads9mo agoHugging Face15Dc-4nderson /gta5_lore-_tipstexttext-generationn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.