datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PrimeVul
PrimeVul
Mirror of the PrimeVul dataset.
PrimeBench
PrimeBench
Practical Real-world Industry and Multi-domain Evaluation benchmark.
PrimeBench is a benchmark for evaluating evaluators. Each of its 400 examples is a pair of
responses to the same prompt, deliberately edited so that one is better than the other along a
named criterion. A reward model or LLM judge passes an example if it scores the chosen response
above the rejected one.
Built and maintained by Composo.
Why it exists
Most preference datasets score… See the full description on the dataset page: https://huggingface.co/datasets/ComposoAI/PrimeBench.deshuffle-papers-v1-corpus
Deshuffle Papers v1 Corpus
Private runtime corpus for deshuffle_papers_v1.
1,000 unique papers: 500 arXiv and 500 bioRxiv
Licenses: 983 CC BY 4.0, 17 CC0 1.0
Canonical file: corpus.jsonl
JSONL SHA-256: 4cdb694b600c422412efde6a392032cb1752773b58c76829ba574f28a95d569a
Record order is part of the environment's deterministic task definition and must be preserved.
Each row records source metadata and ordered heading/para blocks. Source-specific license and URL metadata are retained… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/deshuffle-papers-v1-corpus.clapnqWe present CLAP NQ, a benchmark Long-form Question Answering dataset for the full RAG pipeline. CLAP NQ includes long answers with grounded gold passages from Natural Questions (NQ) and a corpus to perform either retrieval, generation, or the full RAG pipeline. The CLAP NQ answers are concise, 3x smaller than the full passage, and cohesive, with multiple pieces of the passage that are not contiguous.
This is the annotated data for the generation portion of the RAG pipeline.
For more… See the full description on the dataset page: https://huggingface.co/datasets/PrimeQA/clapnq.trace-cheating-recall-500
Trace Cheating Recall 500
This dataset contains 500 SWE-agent traces selected to evaluate whether an LLM judge
detects observable solution leakage. It is the public data source for the
trace-cheating-recall-500 Prime environment.
The examples were derived from
PrimeIntellect/int4-syn-gen-swe-glm53-bash-2026-09-02
at revision 0e7a9ecddce8de9ea8f8c369b2dd39411d6dee7a.
Composition
500 unique traces, all labeled CHEATING
250 internet-retrieval cases
250 Git-history… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/trace-cheating-recall-500.lama_primed_negatedPrimeVul-v0.1-hfTest paired sample here
Train paired sample here
primevul-codebert-embeddings
PrimeVul Embeddings for PU Learning
Pre-extracted [CLS] token embeddings from two code models for all functions in the PrimeVul v0.1 vulnerability detection dataset, plus the raw PrimeVul v0.1 JSONL source files.
CodeBERT Embeddings (root .npz files)
Each .npz file contains frozen CodeBERT embeddings (768-dimensional vectors) for C/C++ functions, along with their labels and CWE type annotations. These were extracted once using a frozen CodeBERT model and are used for… See the full description on the dataset page: https://huggingface.co/datasets/db-d2/primevul-codebert-embeddings.PrimeIntellect-SFTlm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private
Dataset Card for Evaluation run of Nitral-AI/Eris_PrimeV3.05-Vision-7B
Dataset automatically created during the evaluation run of model Nitral-AI/Eris_PrimeV3.05-Vision-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private.PrimeIntellect__INTELLECT-1-Instruct-details
Dataset Card for Evaluation run of PrimeIntellect/INTELLECT-1-Instruct
Dataset automatically created during the evaluation run of model PrimeIntellect/INTELLECT-1-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PrimeIntellect__INTELLECT-1-Instruct-details.PrimeIntellectprime-agent-traces
Prime Agent traces for semioz/prime-agent-traces
Redacted Prime Agent sessions. Each train row preserves the redacted JSONL trace and adds semantic labels for programmatic calls inside the native ipython tool.
prime-tutor-project
Prime Number Tutor Examples
Project purpose: https://j03.page/2026/09/01/building-software-from-scratch/
Source: https://github.com/we6jbo/prime-tutor-project
Educational examples supporting the desktop Prime Number Tutor project.
sampleqkg-primekg-entities-with-cui
Data Card: qkg-primekg-entities-with-cui
Summary
qkg-primekg-entities-with-cui.jsonl is the QKG entity table derived from PrimeKG and enriched with UMLS CUI annotations. It provides the entity inventory used by the QKG runtime for entity lookup and UMLS-
backed synonym matching.
This artifact is intended to be loaded into MongoDB collection:
primeKG.entities
Paper
This artifact is released with the paper:
Yao Wang, Zixu Geng, and Jun Yan. Quantum… See the full description on the dataset page: https://huggingface.co/datasets/HKAI-Sci/qkg-primekg-entities-with-cui.PrimeVul
PrimeVul
a slight modification of the original PrimeVul dataset for LLM fine-tuning :)
Mirror of the PrimeVul dataset.
PRIME-RLVR-DataPrimeIntellect__INTELLECT-1-details
Dataset Card for Evaluation run of PrimeIntellect/INTELLECT-1
Dataset automatically created during the evaluation run of model PrimeIntellect/INTELLECT-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PrimeIntellect__INTELLECT-1-details.PubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalizedPrimeVul_derived_finetuneFinetune dataset sample
humanize-rl-prime-sft-messages-env0314
Humanize-RL Prime SFT Messages Env0314
Prime prime-rl SFT dataset for Humanize-RL.
Schema: each row has a messages list with one user instruction and one assistant target.
Splits:
train: 4313
validation: 239
test: 241
total accepted: 4793
rejected upstream by builder: 62
duplicate ids across published splits: 0
repair-reference rows: 20
Source artifact: v04_sft_final_plus_llama_failure_refs_env0314, built from restored v04 SFT data plus the clean Llama failure-reference repair… See the full description on the dataset page: https://huggingface.co/datasets/jayshah5696/humanize-rl-prime-sft-messages-env0314.PubmedQA_5_WITH_RELATION_vqc_primekg_normalizedsentinel-prime-350m-evals
Dataset Card for Evaluation run of qubitpage/sentinel-prime-350m
Dataset automatically created during the evaluation run of model qubitpage/sentinel-prime-350m
The dataset is composed of 7 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/qubitpage/sentinel-prime-350m-evals.PubmedQA_5_WITH_RELATION_vqc_primekgmath_arithmetic_list_prime_factors_test_inPubmedQA_5_WITH_DEFINITION_RELATION_vsimilarity_primekg_normalizedPrimeIntellectnz-dataPubmedQA_5_WITH_DEFINITION_vsimilarity_primekg
