datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LOREA-cyber-training-data
LOREA-cyber security code-analysis training set
Two corpora live here. The v6_corpus config is the newer one and is what actually trained
LOREA-cyber v6 Pilot. The eight older configs are the v5-era set, kept as-is because they are a
different schema and still useful on their own.
v6_corpus
4,780 train and 151 validation rows in chat format: {"messages": [...], "meta": {...}}, where
messages is a system/user/assistant sequence and meta carries type, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-training-data.Silly-lorebooklore-corpus
COPEAI Lore Corpus
Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus,
mood label definitions, and disclosure copy from
COPEAI — an AI-themed Solana memecoin satire on
Pump.fun.
Compliance frame: Every entry here is fictional in-character satire.
Nothing in this corpus is financial advice, investment guidance, or a
recommendation to transact. COPEAI provides no rights, utility, yield, or
appreciation expectations. The agent names (TRON, CLU, QUORRA, ZUSE… See the full description on the dataset page: https://huggingface.co/datasets/AgentZeroCopeAI/lore-corpus.Ursa-Armored-Core-6-LoreLORE-examples
LORE Examples
A small set of matched multimodal examples from LORE, for the
MIMIC model — enough to try inference,
embedding, and generation across DNA, RNA, and protein modalities without wiring
up your own data.
Each example is a single biological entity (a transcript and/or its protein) with
several co-observed modalities. Rows are drawn from the held-out (validation) split
of MIMIC's training data, so they are in-distribution and length-bounded to the
model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/lorenzo217/HundredCV-Chat.grade-school-math-instructions-Malagasy
Overview
This dataset is a Malagasy adaptation of grade-school-math-instructions.
It consists of arithmetic word problems converted into instruction-answer pairs in Malagasy.
Each entry contains a math problem presented as an instruction, optional contextual input,
and a detailed step-by-step solution in Malagasy.
The dataset is particularly useful for training and evaluating models on arithmetic reasoning and instruction-following tasks in Malagasy, a low-resource language.… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/grade-school-math-instructions-Malagasy.lore-corpus
COPEAI Lore Corpus
Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus,
mood label definitions, and disclosure copy from
COPEAI — an AI-themed Solana memecoin satire on
Pump.fun.
Compliance frame: Every entry here is fictional in-character satire.
Nothing in this corpus is financial advice, investment guidance, or a
recommendation to transact. COPEAI provides no rights, utility, yield, or
appreciation expectations. The agent names (TRON, CLU, QUORRA… See the full description on the dataset page: https://huggingface.co/datasets/Dula23/lore-corpus.LOREA-cyber-eval
LOREA-cyber eval sets
Held-out sets used to benchmark the LOREA-cyber models. Decontaminated 8-gram against the training data,
published so the numbers in the model cards can be reproduced.
These are the sets written for this project. The models are also scored on public benchmarks that aren't
redistributed here: SecQA,
MMLU-Pro,
CyberMetric,
HumanEval.
cyber_mcq (150)
Security knowledge multiple choice across network security, crypto, web/OWASP, malware analysis… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-eval.Kimi-Lorebook-finalskyrim-lore-datasetmsmarco-chunkeval-lorem_ipsum_4000_startwarhammer40k-lore
Warhammer 40K Lore dataset
msmarco-chunkeval-lorem_ipsum_4000_midmsmarco-chunkeval-lorem_ipsum_4000_endmsmarco-chunkeval-lorem_ipsum_400_startmsmarco-chunkeval-lorem_ipsum_500_midPLANE-ood
PLANE Out-of-Distribution Sets
PLANE (phrase-level adjective-noun entailment) is a benchmark to test models on fine-grained compositional inference.
The current dataset contains five sampled splits, used in the supervised experiments of Bertolini et al., 22.
Data Structure
The dataset is organised around five Train/test_split#, each containing a training and test set of circa 60K and 2K.
Features
Each entrance has 6 features: seq, label, Adj_Class, Adj, Nn… See the full description on the dataset page: https://huggingface.co/datasets/lorenzoscottb/PLANE-ood.avos-dark-fantasy-lore-bible
THE DARKNESS STEALS THE LIGHT: Authoritative Lore Bible
Author: Mark Kirkbride
Word Count: ~120000
Canon Status: Master Node v3.0
Raw Full Book Text as Nested Json the-darkness-steals-the-light_Full_Text_v2026 // https://huggingface.co/datasets/TheElim/avos-dark-fantasy-lore-bible/raw/main/the-darkness-steals-the-light_Full_Text_v2026.json
Text Jsonl for Book Training https://huggingface.co/datasets/TheElim/avos-dark-fantasy-lore-bible/raw/main/TDSTL_train.jsonl
I. AI… See the full description on the dataset page: https://huggingface.co/datasets/TheElim/avos-dark-fantasy-lore-bible.Collection-Tononkalo-Malagasy
Collection Tononkalo Malagasy
Dataset Description
This dataset is a collection of Malagasy poems (Tononkalo) scraped from Vetso Serasera. It focuses on creative writing, rhymes, and artistic expression in the Malagasy language.
Important Note: This dataset has been rigorously filtered.
Language Filtering: Poems written primarily in French or English have been removed to ensure a high-quality Malagasy corpus.
Cleaning: Metadata, dates, author signatures inside the text… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/Collection-Tononkalo-Malagasy.msmarco-chunkeval-lorem_ipsum_200_midmsmarco-chunkeval-lorem_ipsum_300_startmsmarco-chunkeval-lorem_ipsum_300_midmsmarco-chunkeval-lorem_ipsum_500_startloremipsum-29k-symbolsbannerlord-lore-dataset
Bannerlord Lore Dataset
Comprehensive lore dataset for Mount & Blade II: Bannerlord - a medieval action RPG by TaleWorlds Entertainment.
Dataset Description
This dataset contains structured lore information extracted from the game, including:
Categories
Category
Description
Languages
Heroes
NPCs, lords, companions
EN, RU, TR
Kingdoms
Major factions (Empire, Battania, etc.)
EN, RU, TR
Settlements
Cities, castles, villages
EN, RU, TR
Cultures… See the full description on the dataset page: https://huggingface.co/datasets/TSEOsiris/bannerlord-lore-dataset.pokemon-lore-instructionsmsmarco-chunkeval-lorem_ipsum_100_startmsmarco-chunkeval-lorem_ipsum_200_endAC6-Lorebook-entries-kimi
