datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text2cypher-2024v1
Neo4j-Text2Cypher (2024) Dataset
The Neo4j-Text2Cypher (2024) Dataset brings together instances from publicly available datasets,
cleaning and organizing them for smoother use. Each entry includes a “question, schema, cypher” triplet at minimum,
with a total of 44,387 instances — 39,554 for training and 4,833 for testing.
An overview of the dataset is shared at Link
Have ideas or insights? Contact us: Neo4j/Team-GenAI
Fields
Fields and their descriptions are as… See the full description on the dataset page: https://huggingface.co/datasets/neo4j/text2cypher-2024v1.Neo-GATE
Dataset card for Neo-GATE
Homepage: https://mt.fbk.eu/neo-gate/
Dataset summary
Neo-GATE is a bilingual corpus designed to benchmark the ability of machine translation (MT) systems to translate from English into Italian using gender-inclusive neomorphemes.
It is built upon GATE (Rarrick et al., 2023), a benchmark for the evaluation of gender rewriters and gender bias in MT.
Neo-GATE includes 841 test entries (Neo-GATE.tsv) and 100 dev entries (Neo-GATE-dev.tsv).
Each… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Neo-GATE.smolgpt-markdown-stories
SmolGPT-Fables Stories
A deterministic, text-only corpus of 96,000 original English
Markdown stories built for SmolGPT-Fables. Every row is one complete supervised
story example with an exact prompt / completion boundary, a requested scene
count from one to six, and plain-language conditioning fields.
No model, API, browser, or network service was used to create this dataset.
Dataset summary
96,000 stories across 96,000 isolated story families
25 genres and all… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/smolgpt-markdown-stories.LIT-RAGBench
LIT-RAGBench
LIT-RAGBench is a benchmark for evaluating generator capabilities in Retrieval-Augmented Generation (RAG). It focuses on whether a model can answer questions correctly given retrieved documents, independent of retrieval quality. The benchmark covers five categories: Integration, Reasoning, Logic, Table, and Abstention.
Dataset Summary
LIT-RAGBench contains:
114 human-constructed Japanese questions
An English version generated by machine translation with… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/LIT-RAGBench.omnimcp_graphrag_neo4j_cypher_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_neo4j_cypher_teaser.Nemotron-Personas-Korea
Nemotron-Personas-Korea
우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템
A compound AI approach to personas grounded in real-world distributions
데이터셋 개요 (Overview)
Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.
Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/neomoon50/Nemotron-Personas-Korea.korean-llm-citation-baseline-2026
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018479 (Zenodo record).
Cite as:
@dataset{neogenesis_20018479,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Korean LLM Citation Baseline 2026 (Neo Genesis GEO Measurement)},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018479},
url = {https://doi.org/10.5281/zenodo.20018479}
}
Korean LLM Citation Baseline 2026 (Neo Genesis GEO… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/korean-llm-citation-baseline-2026.cross-agent-review-queue-2026
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018477 (Zenodo record).
Cite as:
@dataset{neogenesis_20018477,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Cross-Agent Code Review Queue (Codex <-> Claude, Neo Genesis 2026)},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018477},
url = {https://doi.org/10.5281/zenodo.20018477}
}
Cross-Agent Code Review Queue (Codex <-> Claude, Neo… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/cross-agent-review-queue-2026.MELD-MPCA
Dataset Information
the dataset is a jsonl file containing each dialogue (context) per line.
Field
Amount
Dialogue (context/line)
1022
Diff User
260
each dialogue context messages of a conversation, with those informations:
user
content
emotion
type
n_turn
summary
traits
distanglement
ref_speaker
ref_utterance
tar_speaker
selected_speaker
The user of the message
The content of the message
The emotion of the user
Either a positive or negative emotion
The… See the full description on the dataset page: https://huggingface.co/datasets/neoluigi/MELD-MPCA.neo1.0-benchmark
Neo — LLM-Halluzinations-Benchmark (Deutsch)
134 Fragen (Multiple Choice + offene Fragen) in Deutsch, gebaut, um typische
Halluzinationen kleiner lokaler LLMs zu provozieren — nicht um Weltwissen abzufragen.
Genutzt für das eigene Fine-Tune Dimitrex93/neo1.0-3b
und als SFT-Ground-Truth.
English: 134 German-language questions (MC + open) designed to trigger the hallucination
patterns of small local LLMs. Used as benchmark and SFT ground truth for
neo1.0-3b. Evaluation harness:… See the full description on the dataset page: https://huggingface.co/datasets/Dimitrex93/neo1.0-benchmark.neo_sft_phase2_conversations
1. The original dataset can be found at:
https://huggingface.co/datasets/m-a-p/neo_sft_phase2
2. Split multi-turn conversations into individual single-turn samples
Approach: Treat each round of dialogue as an independent question-and-answer pair, and construct the sample using contextual information.
Specific operations:
For each "conversations", iterate through each round of dialogue.
Concatenate the "value" of the current "human" round with the dialogue from all… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_conversations.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.neo_ara_v2sora-multi-device-orchestration-2026
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018481 (Zenodo record).
Cite as:
@dataset{neogenesis_20018481,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Sora Multi-Device Orchestration Architecture 2026},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018481},
url = {https://doi.org/10.5281/zenodo.20018481}
}
Sora Multi-Device Orchestration Architecture 2026
A reference… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/sora-multi-device-orchestration-2026.neo-sql-reasoning-combined
Neo SQL + Reasoning Combined Dataset
Combined SFT dataset for fine-tuning SQL, reasoning, and math models.
Built for the neo-deep-agent-lab project.
Sources & Proportions
Source
Proportion
Records
Focus
gretelai/synthetic_text_to_sql
50%
~5,000
SQL generation
nohurry/Opus-4.6-Reasoning-3000x-filtered
30%
~2,100
Reasoning
openai/gsm8k
20%
~1,400
Math problems
Format
All samples are normalized to SFT chat format:
{
"messages": [… See the full description on the dataset page: https://huggingface.co/datasets/Shumatsurontek/neo-sql-reasoning-combined.neolurk-dataset
Neolurk.org Memes Dataset
Этот датасет содержит очищенный текстовый корпус русской интернет-энциклопедии мемов Neolurk (современный преемник Lurkmore).
Описание
Данные предназначены для обучения и файнтюнинга больших языковых моделей (LLM), помогая им лучше понимать русскоязычный интернет-фольклор, сленг, мемы и исторический контекст субкультур.
Количество статей: 48 234 статьи
Формат: JSON Lines (.jsonl), где каждая строка содержит:
pageid (int): уникальный… See the full description on the dataset page: https://huggingface.co/datasets/Alex01837178373/neolurk-dataset.H4rmony_dpo
Citation Information
@article{neovalle2024h4rmony,
author = {Vallego, Jorge},
title = {H4rmony DPO Dataset},
howpublished = {Hugging Face Hub},
year = {2024},
url = {https://huggingface.co/datasets/neovalle/H4rmony_dpo}
}
This dataset is based on neovalle/H4rmony, and optimised to the format required by DPOTrainer from the trl library.
whylab-gemini-2-5-docker-validation
🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced.
DOI
This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.quant-v11-ensemble-6alpha-specs-2026
DOI
This dataset is citable via DataCite DOI 10.5281/zenodo.20018487 (Zenodo record).
Cite as:
@dataset{neogenesis_20018487,
author = {Heo, Yesol and Neo Genesis Lab},
title = {Quant v11 Ensemble 6-Alpha Specs & Risk Engineering 2026},
year = 2026,
publisher = {Zenodo},
doi = {10.5281/zenodo.20018487},
url = {https://doi.org/10.5281/zenodo.20018487}
}
Quant v11 Ensemble - 6-Alpha Specs & Risk Engineering 2026
A… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/quant-v11-ensemble-6alpha-specs-2026.neo_sft_phase2_single
dataset
The original dataset can be found at: https://huggingface.co/datasets/m-a-p/neo_sft_phase2
Use the following code to select two-turn conversations for your SFT dataset.
code
import json
def process_conversations(input_file, output_file):
with open(input_file, 'r', encoding='utf-8') as f_in, \
open(output_file, 'w', encoding='utf-8') as f_out:
data = json.load(f_in)
for item in data:
conversations =… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_single.research-papers-gpt-neox
abhi26/research-papers-gpt-neox
This dataset contains processed research papers optimized for GPT-NeoX-20B training.
The text has been cleaned, chunked to 2048 tokens, and formatted for causal language modeling.
Dataset Details
Total Samples: 9993
Unique Papers: 1017
Average Tokens per Sample: 1965.4
Token Range: 10 - 91659
Max Token Limit: 2048
Source Subdirectories: 1
Dataset Structure
Each sample contains:
text: The processed research paper text or chunk… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/research-papers-gpt-neox.neo_sft_phase2_multi
1. The original dataset can be found at:
https://huggingface.co/datasets/m-a-p/neo_sft_phase2
2. Split multi-turn conversations into individual single-turn samples
Approach: Treat each round of dialogue as a separate question-and-answer pair, and construct the sample by leveraging the contextual information.
Specific Operations:
For each "conversation," iterate through all the dialogue rounds.
Concatenate the "value" of all "human" turns within each "conversation" to… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_multi.neo_ara_v1pile-neox-uint16-partsTokenized uint16 shard parts for language-model pretraining.
Original source: The Pile / NeoX-style preprocessing.
neocortirrhea-lexicon
Dataset Card: Neocortirrhea Lexicon Entry
Summary
This dataset entry defines and contextualizes the psychological, neurological, and somatic neologism Neocortirrhea.
Dataset Structure
JSON Lines Representation (data.jsonl)
{
"term": "Neocortirrhea",
"part_of_speech": "noun",
"phonetic": "/ˌniː.oʊˌkɔːr.tɪˈriː.ə/",
"etymology": "Neocortex (higher-order cognitive processing) + -rrhea (Greek rhoia: abnormal/excessive flow or… See the full description on the dataset page: https://huggingface.co/datasets/dlewicki/neocortirrhea-lexicon.Moon-1-DataMoon-2-Data
