datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
juliet_test_suite_c_1_3
Dataset Card for the Juliet Test Suite 1.3
Dataset Summary
This Datasets contains all test cases from the NIST's Juliet test suite for the C and C++ programming languages. The dataset contains a benign and a defective implementation of each sample, which have been extracting by means of the OMITGOOD and OMITBAD preprocessor macros of the Juliet test suite.
Supported Tasks and Leaderboards
Software defect prediction, code clone detection.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/LorenzH/juliet_test_suite_c_1_3.japan-procurement-open-data
Japan Public Procurement & Company Open Data
Machine-readable extracts of Japanese public-procurement and company open data, compiled and normalised by LoreaTec for japan-tenders.loreatec.jp and bizsearch.loreatec.jp. Everything here comes from official Japanese government sources; the value added is the cleaning, joining and the derived analysis (contract series and re-tender predictions).
Updated monthly. The authoritative, always-current copy is… See the full description on the dataset page: https://huggingface.co/datasets/loreatec/japan-procurement-open-data.image-splicing-deepfake-mix-newtrivia_qa
Dataset Card for "trivia_qa"
Dataset Summary
TriviaqQA is a reading comprehension dataset containing over 650K
question-answer-evidence triples. TriviaqQA includes 95K question-answer
pairs authored by trivia enthusiasts and independently gathered evidence
documents, six per question on average, that provide high quality distant
supervision for answering the questions.
Supported Tasks and Leaderboards
More Information Needed
Languages… See the full description on the dataset page: https://huggingface.co/datasets/lorenzofalappa/trivia_qa.LOREA-cyber-training-data
LOREA-cyber security code-analysis training set
Two corpora live here. The v6_corpus config is the newer one and is what actually trained
LOREA-cyber v6 Pilot. The eight older configs are the v5-era set, kept as-is because they are a
different schema and still useful on their own.
v6_corpus
4,780 train and 151 validation rows in chat format: {"messages": [...], "meta": {...}}, where
messages is a system/user/assistant sequence and meta carries type, domain, and… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-training-data.vaovao_malagasy_sentiment_corpus
Dataset Card for Vaovao Malagasy Sentiment Corpus (VMSC)
Dataset Summary
The Vaovao Malagasy Sentiment Corpus (VMSC) is the first publicly available, manually annotated sentiment analysis dataset for the Malagasy language (mg). It contains 5,041 sentences extracted from news articles (vaovao) published between 2022 and 2023.
Each sentence is labeled with binary sentiment (Positive or Negative). The dataset was created to address the scarcity of resources for Malagasy NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/vaovao_malagasy_sentiment_corpus.Silly-lorebookimage-splicing-deepfake-mixlore-corpus
COPEAI Lore Corpus
Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus,
mood label definitions, and disclosure copy from
COPEAI — an AI-themed Solana memecoin satire on
Pump.fun.
Compliance frame: Every entry here is fictional in-character satire.
Nothing in this corpus is financial advice, investment guidance, or a
recommendation to transact. COPEAI provides no rights, utility, yield, or
appreciation expectations. The agent names (TRON, CLU, QUORRA, ZUSE… See the full description on the dataset page: https://huggingface.co/datasets/AgentZeroCopeAI/lore-corpus.Ursa-Armored-Core-6-LoreLORE-examples
LORE Examples
A small set of matched multimodal examples from LORE, for the
MIMIC model — enough to try inference,
embedding, and generation across DNA, RNA, and protein modalities without wiring
up your own data.
Each example is a single biological entity (a transcript and/or its protein) with
several co-observed modalities. Rows are drawn from the held-out (validation) split
of MIMIC's training data, so they are in-distribution and length-bounded to the
model's context… See the full description on the dataset page: https://huggingface.co/datasets/polymathic-ai/LORE-examples.multi-target-spacecraft-pose-estimation
Multi-target Synthetic Dataset for Spacecraft Pose Estimation
Overview
This dataset was developed for 6D pose estimation of unseen, non-cooperative spacecraft in proximity-operations scenarios. Most existing datasets focus on a single target, which leads models to overfit to a specific spacecraft and limits their ability to generalize to previously unseen targets.
To address this limitation, the present dataset is multi-target and includes a wide variety of spacecraft… See the full description on the dataset page: https://huggingface.co/datasets/lorenzobottelli/multi-target-spacecraft-pose-estimation.loan-approval-dataset
Loan Approval Dataset
Dataset for experimenting with binary classification models for
loan approval prediction.
Dataset structure
The dataset contains two splits:
train: 614 records
test: 367 records
The training dataset contains the target variable:
Loan_Status
where:
Y = Loan approved
N = Loan rejected
Features
Loan_ID
Gender
Married
Dependents
Education
Self_Employed
ApplicantIncome
CoapplicantIncome
LoanAmount
Loan_Amount_Term… See the full description on the dataset page: https://huggingface.co/datasets/LoreRodriguez/loan-approval-dataset.HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/lorenzo217/HundredCV-Chat.E2MoCase
E2MoCase: summary
E2MoCase is a novel curated dataset linking news stories about real-world legal cases to (i) the concrete events they describe, (ii) the emotions they evoke, and (iii) the moral foundations they frame. Articles are segmented into paragraphs, and each paragraph is independently annotated with aligned event (triggering words and involved entities), emotion labels, and moral labels, giving researchers a fine-grained lens on narrative bias. The resource paper… See the full description on the dataset page: https://huggingface.co/datasets/lorenzozan/E2MoCase.grade-school-math-instructions-Malagasy
Overview
This dataset is a Malagasy adaptation of grade-school-math-instructions.
It consists of arithmetic word problems converted into instruction-answer pairs in Malagasy.
Each entry contains a math problem presented as an instruction, optional contextual input,
and a detailed step-by-step solution in Malagasy.
The dataset is particularly useful for training and evaluating models on arithmetic reasoning and instruction-following tasks in Malagasy, a low-resource language.… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/grade-school-math-instructions-Malagasy.lore-corpus
COPEAI Lore Corpus
Open dataset of in-character lore, agent dossiers, blog dispatches, FAQ corpus,
mood label definitions, and disclosure copy from
COPEAI — an AI-themed Solana memecoin satire on
Pump.fun.
Compliance frame: Every entry here is fictional in-character satire.
Nothing in this corpus is financial advice, investment guidance, or a
recommendation to transact. COPEAI provides no rights, utility, yield, or
appreciation expectations. The agent names (TRON, CLU, QUORRA… See the full description on the dataset page: https://huggingface.co/datasets/Dula23/lore-corpus.LOREA-cyber-eval
LOREA-cyber eval sets
Held-out sets used to benchmark the LOREA-cyber models. Decontaminated 8-gram against the training data,
published so the numbers in the model cards can be reproduced.
These are the sets written for this project. The models are also scored on public benchmarks that aren't
redistributed here: SecQA,
MMLU-Pro,
CyberMetric,
HumanEval.
cyber_mcq (150)
Security knowledge multiple choice across network security, crypto, web/OWASP, malware analysis… See the full description on the dataset page: https://huggingface.co/datasets/MK4-Research/LOREA-cyber-eval.Kimi-Lorebook-finalskyrim-lore-datasetmsmarco-chunkeval-lorem_ipsum_4000_startcorpus-en-es
Dataset Card for "corpus-en-es"
More Information needed
degeneration-probe-instruct-token-level-balanced
degeneration-probe-instruct-token-level-balanced
Downsampled (1:3 positive:negative) variant of luca-sartori/degeneration-probe-instruct-token-level.
An example is considered positive if it contains at least one token with repetition >= 0.8 in the chunk_summary field. The downsampling keeps all positive examples and adds a random subset of negative examples in a 1:3 ratio.
Examples whose chunk_summary contained no scored tokens (every repetition value null) have been dropped, since… See the full description on the dataset page: https://huggingface.co/datasets/lorenzo0312/degeneration-probe-instruct-token-level-balanced.ru-ky-synthetic-loresmt2026
Russian-Kyrgyz Synthetic Parallel Corpus (LoResMT 2026)
A synthetic parallel corpus for Russian-Kyrgyz machine translation, created for the LoResMT 2026 Turkic Languages Translation Challenge.
Dataset Description
This dataset contains synthetic Russian-Kyrgyz parallel sentences generated by translating:
FineWeb-2 Kyrgyz data → Russian (back-translation using Gemma3-27B and Qwen3-235B)
SiberianPersonaChat dialogues → Kyrgyz (translation using GPT-4o)
All data has been… See the full description on the dataset page: https://huggingface.co/datasets/Novokshanov/ru-ky-synthetic-loresmt2026.warhammer40k-lore
Warhammer 40K Lore dataset
malagasy-sentence
Overview
This dataset consists of clean, structured sentences extracted via Optical Character Recognition (OCR) from approximately 1GB of Malagasy thesis documents. These documents were collected based on educational, cultural, and linguistic themes.
The dataset is saved in CSV format, and is particularly useful for NLP tasks involving sentence-level modeling in Malagasy — a low-resource language.
Dataset Details
Language: Malagasy
Source: OCR'd academic thesis… See the full description on the dataset page: https://huggingface.co/datasets/Lo-Renz-O/malagasy-sentence.msmarco-chunkeval-lorem_ipsum_4000_midmsmarco-chunkeval-lorem_ipsum_4000_endla-usc_valbadia_loresmt24
Dataset Card: Ladin (Val Badia) - Monolingual (La Usc di Ladins)
Overview
Source Paper: "Rule-Based, Neural and LLM Back-Translation: Comparative Insights from a Variant of Ladin"
Description:
This dataset contains monolingual sentences in Ladin (Val Badia) from the newspaper "La Usc di Ladins" (https://www.lausc.it), which has been archived since 2008. The newspaper offers texts in five variants of Ladin, corresponding to the five Ladin valleys.
We extracted 1,937,608… See the full description on the dataset page: https://huggingface.co/datasets/sfrontull/la-usc_valbadia_loresmt24.msmarco-chunkeval-lorem_ipsum_400_start
