datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lm-eval-results-Kukedlc-NeoCortex-7B-slerp-private
Dataset Card for Evaluation run of Kukedlc/NeoCortex-7B-slerp
Dataset automatically created during the evaluation run of model Kukedlc/NeoCortex-7B-slerp
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-NeoCortex-7B-slerp-private.NIAH-gpt-neox-20bovos-wake-word-bench-synthetic-wakewords-hey_neon
OVOS wake_word bench — synthetic-wakewords-hey_neon
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_neon.EleutherAI__gpt-neox-20b-details
Dataset Card for Evaluation run of EleutherAI/gpt-neox-20b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neox-20b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neox-20b-details.TrapQA
TrapQA
TrapQA is a benchmark suite for evaluating whether language models remain faithful to decisive constraints when misleading priors or salient associations point toward an incorrect answer.
It contains two complementary subsets:
ScientistQA: entity-level factual disambiguation between two candidate scientists.
Real-Life Constrained QA: everyday two-option scenarios where physical, spatial, procedural, or medium-specific constraints conflict with intuitive shortcuts.
All… See the full description on the dataset page: https://huggingface.co/datasets/NeoHugh/TrapQA.togethercomputer__GPT-NeoXT-Chat-Base-20B-details
Dataset Card for Evaluation run of togethercomputer/GPT-NeoXT-Chat-Base-20B
Dataset automatically created during the evaluation run of model togethercomputer/GPT-NeoXT-Chat-Base-20B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-NeoXT-Chat-Base-20B-details.neopolita__jessi-v0.5-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.5-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.5-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.5-falcon3-7b-instruct-details.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.EleutherAI__gpt-neo-125m-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-125m
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-125m
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-125m-details.repro-skill-neologisms-traces
Agent traces
Agent sessions published from a Trackio Logbook.
EleutherAI__gpt-neo-2.7B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-2.7B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-2.7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-2.7B-details.open-neo__Kyro-n1-7B-details
Dataset Card for Evaluation run of open-neo/Kyro-n1-7B
Dataset automatically created during the evaluation run of model open-neo/Kyro-n1-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/open-neo__Kyro-n1-7B-details.EleutherAI__gpt-neo-1.3B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-1.3B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-1.3B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-1.3B-details.open-neo__Kyro-n1-3B-details
Dataset Card for Evaluation run of open-neo/Kyro-n1-3B
Dataset automatically created during the evaluation run of model open-neo/Kyro-n1-3B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/open-neo__Kyro-n1-3B-details.neopets-packai-brand-mention-baseline-2026
AI Brand Mention Baseline 2026
A longitudinal benchmark dataset measuring how frontier LLMs (Gemini 2.5,
GPT-4 class, Claude class) mention a single AI-native company (Neo
Genesis) when prompted with content-gap probes. First open dataset of
its kind for GEO (Generative Engine Optimization) research.
Metric
Value
Measurements
486
Window
2026-04-28 to 2026-05-07 (10 days)
Distinct seed prompts
30
Categories
6 (definition, pricing, comparison, problem_solving… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/ai-brand-mention-baseline-2026.BlackBeenie__Neos-Llama-3.1-8B-details
Dataset Card for Evaluation run of BlackBeenie/Neos-Llama-3.1-8B
Dataset automatically created during the evaluation run of model BlackBeenie/Neos-Llama-3.1-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BlackBeenie__Neos-Llama-3.1-8B-details.BlackBeenie__Neos-Phi-3-14B-v0.1-details
Dataset Card for Evaluation run of BlackBeenie/Neos-Phi-3-14B-v0.1
Dataset automatically created during the evaluation run of model BlackBeenie/Neos-Phi-3-14B-v0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BlackBeenie__Neos-Phi-3-14B-v0.1-details.Kimargin__GPT-NEO-1.3B-wiki-details
Dataset Card for Evaluation run of Kimargin/GPT-NEO-1.3B-wiki
Dataset automatically created during the evaluation run of model Kimargin/GPT-NEO-1.3B-wiki
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Kimargin__GPT-NEO-1.3B-wiki-details.BlackBeenie__Neos-Gemma-2-9b-details
Dataset Card for Evaluation run of BlackBeenie/Neos-Gemma-2-9b
Dataset automatically created during the evaluation run of model BlackBeenie/Neos-Gemma-2-9b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/BlackBeenie__Neos-Gemma-2-9b-details.neopolita__jessi-v0.6-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.6-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.6-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.6-falcon3-7b-instruct-details.neopolita__jessi-v0.3-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.3-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.3-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.3-falcon3-7b-instruct-details.bamec66557__MISCHIEVOUS-12B-Mix_Neo-details
Dataset Card for Evaluation run of bamec66557/MISCHIEVOUS-12B-Mix_Neo
Dataset automatically created during the evaluation run of model bamec66557/MISCHIEVOUS-12B-Mix_Neo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bamec66557__MISCHIEVOUS-12B-Mix_Neo-details.neopolita__jessi-v0.1-qwen2.5-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.1-qwen2.5-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.1-qwen2.5-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.1-qwen2.5-7b-instruct-details.neopolita__jessi-v0.2-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.2-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.2-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.2-falcon3-7b-instruct-details.neopolita__jessi-v0.1-virtuoso-small-details
Dataset Card for Evaluation run of neopolita/jessi-v0.1-virtuoso-small
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.1-virtuoso-small
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.1-virtuoso-small-details.neopolita__jessi-v0.4-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.4-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.4-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.4-falcon3-7b-instruct-details.neopolita__jessi-v0.1-falcon3-10b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.1-falcon3-10b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.1-falcon3-10b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.1-falcon3-10b-instruct-details.neopolita__jessi-v0.2-falcon3-10b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.2-falcon3-10b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.2-falcon3-10b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.2-falcon3-10b-instruct-details.allura-org__TQ2.5-14B-Neon-v1-details
Dataset Card for Evaluation run of allura-org/TQ2.5-14B-Neon-v1
Dataset automatically created during the evaluation run of model allura-org/TQ2.5-14B-Neon-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/allura-org__TQ2.5-14B-Neon-v1-details.
