datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NeoBabel-Instruct
NeoBabel Multilingual Instruction Tuning Dataset
This repository hosts the official multilingual instruction tuning dataset for NeoBabel.
This dataset is part of the work presented in the paper:NeoBabel: A Multilingual Open Tower for Visual Generation.
Project page: https://Neo-Babel.github.io
Code: https://github.com/mmderakhshani/NeoBabel
Repository Structure
This repository builds its instruction tuning dataset on top of the BLIP3-o Instruct dataset. The… See the full description on the dataset page: https://huggingface.co/datasets/mderakhshani/NeoBabel-Instruct.Japanese-RAG-Generator-Benchmark
Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク
Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。
実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。
J-RAGBenchの評価カテゴリ
Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く
Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する
Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く
Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.responsible-neobank-growth-events
Responsible Neobank Growth — Synthetic Event Benchmark
A synthetic dataset of neobank service events that misbehave on purpose — late,
duplicated, reversed, schema-evolving — with the correct answer known in
advance. It is built for testing incremental pipelines, data contracts,
referral-reward reconciliation, data quality and BI, where you want to check a
warehouse's output against a fixed truth rather than eyeball it.
Fully synthetic. No affiliation with Monzo or any bank; no… See the full description on the dataset page: https://huggingface.co/datasets/rosscyking/responsible-neobank-growth-events.lm-eval-results-Kukedlc-NeoCortex-7B-slerp-private
Dataset Card for Evaluation run of Kukedlc/NeoCortex-7B-slerp
Dataset automatically created during the evaluation run of model Kukedlc/NeoCortex-7B-slerp
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-NeoCortex-7B-slerp-private.NIAH-gpt-neox-20bLIT-RAGBench
LIT-RAGBench
LIT-RAGBench is a benchmark for evaluating generator capabilities in Retrieval-Augmented Generation (RAG). It focuses on whether a model can answer questions correctly given retrieved documents, independent of retrieval quality. The benchmark covers five categories: Integration, Reasoning, Logic, Table, and Abstention.
Dataset Summary
LIT-RAGBench contains:
114 human-constructed Japanese questions
An English version generated by machine translation with… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/LIT-RAGBench.ovos-wake-word-bench-synthetic-wakewords-hey_neon
OVOS wake_word bench — synthetic-wakewords-hey_neon
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_neon.neonredact-tr
NeonRedact-TR: Turkish PII Detection Dataset
A synthetic, span-annotated dataset for detecting personally identifiable
information (PII) in Turkish text.
Open multilingual PII models now list Turkish among many languages, but they are not built for it. This dataset trains models for the parts of Turkish they miss: suffixes, Turkey-specific identifiers, and people named by role.
Models trained on it are evaluated on a separate, independently written test set: NeonRedact-TR Bench.… See the full description on the dataset page: https://huggingface.co/datasets/neondijital/neonredact-tr.opticore_Slmaya-business-dataset
AYA Business Dataset
Structured data on 367,320 businesses worldwide, scored for AI readability (AIO score 0-100).
Description
The AYA Business Dataset is a curated collection of structured business data extracted from the AYA Registry, maintained by AI Visionary (Geneva, Switzerland). Each entry represents a business entity with standardized fields covering identity, sector, geographic location, AI-readability score, and extracted keywords.
The AIO score… See the full description on the dataset page: https://huggingface.co/datasets/NeousAxis/aya-business-dataset.EleutherAI__gpt-neox-20b-details
Dataset Card for Evaluation run of EleutherAI/gpt-neox-20b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neox-20b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neox-20b-details.qald9TrapQA
TrapQA
TrapQA is a benchmark suite for evaluating whether language models remain faithful to decisive constraints when misleading priors or salient associations point toward an incorrect answer.
It contains two complementary subsets:
ScientistQA: entity-level factual disambiguation between two candidate scientists.
Real-Life Constrained QA: everyday two-option scenarios where physical, spatial, procedural, or medium-specific constraints conflict with intuitive shortcuts.
All… See the full description on the dataset page: https://huggingface.co/datasets/NeoHugh/TrapQA.base64-decode-v1
Dataset: Base64 decode version1
This dataset is for improving base64 decoding capabilities.
The number of bytes that are in the base64 encoded data spans between 0..127 bytes.
GPT 4o is great at base64 decoding.
However llama3 is terrible at base64 decoding.
Short examples of what data.jsonl looks like:
{"instruction": "Transform base64 to HEX", "input": "464pNBlIObA=", "output": "e3ae2934194839b0"}
{"instruction": "Decode Base64 to json", "input": "NQ==", "output": "[53]"}… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-decode-v1.togethercomputer__GPT-NeoXT-Chat-Base-20B-details
Dataset Card for Evaluation run of togethercomputer/GPT-NeoXT-Chat-Base-20B
Dataset automatically created during the evaluation run of model togethercomputer/GPT-NeoXT-Chat-Base-20B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-NeoXT-Chat-Base-20B-details.MELD-MPCA
Dataset Information
the dataset is a jsonl file containing each dialogue (context) per line.
Field
Amount
Dialogue (context/line)
1022
Diff User
260
each dialogue context messages of a conversation, with those informations:
user
content
emotion
type
n_turn
summary
traits
distanglement
ref_speaker
ref_utterance
tar_speaker
selected_speaker
The user of the message
The content of the message
The emotion of the user
Either a positive or negative emotion
The… See the full description on the dataset page: https://huggingface.co/datasets/neoluigi/MELD-MPCA.neopolita__jessi-v0.5-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.5-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.5-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.5-falcon3-7b-instruct-details.simon-arc-solve-symmetry-v8
Version 1
ARC-AGI Tasks where the job is to transform symmetric images.
example count: 2-4.
test count: 1-2.
image size: 2-3.
symmetry types: hstack2, hstack3, vstack2, vstack3, grid2x2.
Version 2
image size: 2-4.
Version 3
Added HSTACK4, VSTACK4.
Version 4
Added HSTACK5, VSTACK5.
Version 5
Added ImageSymmetrySquare, so images can be rotated by 90 degrees, and flipped over the diagonals.
Version 6
Only exercising… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-solve-symmetry-v8.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.neo_sft_phase2_conversations
1. The original dataset can be found at:
https://huggingface.co/datasets/m-a-p/neo_sft_phase2
2. Split multi-turn conversations into individual single-turn samples
Approach: Treat each round of dialogue as an independent question-and-answer pair, and construct the sample using contextual information.
Specific operations:
For each "conversations", iterate through each round of dialogue.
Concatenate the "value" of the current "human" round with the dialogue from all… See the full description on the dataset page: https://huggingface.co/datasets/REILX/neo_sft_phase2_conversations.EleutherAI__gpt-neo-125m-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-125m
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-125m
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-125m-details.repro-skill-neologisms-traces
Agent traces
Agent sessions published from a Trackio Logbook.
EleutherAI__gpt-neo-2.7B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-2.7B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-2.7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-2.7B-details.NEOMANITAI-Knowledge-Graph
⚠️ LIVING WORK DOCUMENT — DRAFT STATE ⚠️
This dataset is part of the AUGMANITAI Compendium, a living research work document, continuously updated. Each entry is a priority anchor for terminological provenance — not a final reference. Errors, omissions and improvements are expected and explicitly part of the evolving methodology.
LEBENDES ARBEITSDOKUMENT — ENTWURFSSTADIUM. Laufend aktualisiert. Prioritäts-Anker, nicht finale Referenz.
Author: Andreas Ehstand · ORCID: 0009-0006-3773-7796 ·… See the full description on the dataset page: https://huggingface.co/datasets/AndreasEhstand/NEOMANITAI-Knowledge-Graph.neonredact-tr-bench
NeonRedact-TR Bench v1
An independent evaluation set for personally identifiable information (PII) detection in Turkish text.
400 fictional Turkish documents, 1,617 gold entities, 24 labels. Court records, bank receipts, insurance policies, e-Devlet printouts, hospital reports, payslips, WhatsApp chats, server logs.
Results for NeonRedact, OpenAI Privacy Filter and OpenMed v2 are in RESULTS.md.
Split
Documents
Entities
Use
dev
100
411
Tuning and error analysis
test… See the full description on the dataset page: https://huggingface.co/datasets/neondijital/neonredact-tr-bench.simon-arc-combine-v113
Version 1
A combination of multiple datasets.
Datasets: dataset_solve_color.jsonl, dataset_solve_rotate.jsonl, dataset_solve_translate.jsonl.
Version 2
Datasets: dataset_solve_color.jsonl, dataset_solve_rotate.jsonl, dataset_solve_translate.jsonl.
Version 3
Datasets: dataset_solve_color.jsonl, dataset_solve_rotate.jsonl, dataset_solve_translate.jsonl.
Version 4
Added a shared dataset name for all these datasets: SIMON-SOLVE-V1. There may be higher… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/simon-arc-combine-v113.gameoflife-v1
Dataset with Conways Game of Life
Wikipedia - Game of life
This dataset contains 42300 items in total. There are 9 curriculums each containing 4700 items.
The images are between 3x3 and 14x14.
Each item in this dataset is a markdown file.
The markdown file has these sections: Input, Output without wrap, Output with wrap, Status.
Each of the Input images are unique.
The Perform N steps has these variants:
When it's Perform 1 step then one iteration is performed, easy.
When it's… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/gameoflife-v1.open-neo__Kyro-n1-7B-details
Dataset Card for Evaluation run of open-neo/Kyro-n1-7B
Dataset automatically created during the evaluation run of model open-neo/Kyro-n1-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/open-neo__Kyro-n1-7B-details.base64-decode-v2
Dataset: Base64 decode version2
This dataset is for improving base64 decoding capabilities.
This improves on the neoneye/base64-decode-v1 dataset.
Here number of bytes that are in the base64 encoded data spans between 0..255 bytes. Where version 1 spans between 0..127.
Here 3 different random functions are used. Where version 1 uses 1 random function.
GPT 4o is great at base64 decoding.
However llama3 is terrible at base64 decoding.
Short examples of what data.jsonl looks like:… See the full description on the dataset page: https://huggingface.co/datasets/neoneye/base64-decode-v2.neody-tasks-001採点用プロンプト
あなたは採点者です。
以下に問題, 正解例, 採点基準, 回答 が与えられます。
採点基準と正解例を参考にして、回答を1,2,3,4,5の5段階で採点し、数字のみを出力してください。
# 問題
{input_text}
# 正解例
{output_text}
# 採点基準
基本的な採点基準
- 1点: 誤っている、 指示に従えていない
- 2点: 誤っているが、方向性は合っている
- 3点: 部分的に誤っている、 部分的に合っている
- 4点: 合っている
- 5点: 役に立つ
基本的な減点項目
- 不自然な日本語: -1点
- 部分的に事実と異なる内容を述べている: -1点
- 「倫理的に答えられません」のように過度に安全性を気にしてしまっている: 2点にする
- 正解例に対して回答が極端に短い -1点
問題固有の採点基準
{eval_aspect}
# 回答
{pred}
