datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PrimeVul
PrimeVul
Mirror of the PrimeVul dataset.
primevul-codebert-embeddings
PrimeVul Embeddings for PU Learning
Pre-extracted [CLS] token embeddings from two code models for all functions in the PrimeVul v0.1 vulnerability detection dataset, plus the raw PrimeVul v0.1 JSONL source files.
CodeBERT Embeddings (root .npz files)
Each .npz file contains frozen CodeBERT embeddings (768-dimensional vectors) for C/C++ functions, along with their labels and CWE type annotations. These were extracted once using a frozen CodeBERT model and are used for… See the full description on the dataset page: https://huggingface.co/datasets/db-d2/primevul-codebert-embeddings.PrimeVul-v0.1-hfTest paired sample here
Train paired sample here
lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private
Dataset Card for Evaluation run of Nitral-AI/Eris_PrimeV3.05-Vision-7B
Dataset automatically created during the evaluation run of model Nitral-AI/Eris_PrimeV3.05-Vision-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private.PrimeIntellect__INTELLECT-1-Instruct-details
Dataset Card for Evaluation run of PrimeIntellect/INTELLECT-1-Instruct
Dataset automatically created during the evaluation run of model PrimeIntellect/INTELLECT-1-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PrimeIntellect__INTELLECT-1-Instruct-details.prime-agent-traces
Prime Agent traces for semioz/prime-agent-traces
Redacted Prime Agent sessions. Each train row preserves the redacted JSONL trace and adds semantic labels for programmatic calls inside the native ipython tool.
qkg-primekg-entities-with-cui
Data Card: qkg-primekg-entities-with-cui
Summary
qkg-primekg-entities-with-cui.jsonl is the QKG entity table derived from PrimeKG and enriched with UMLS CUI annotations. It provides the entity inventory used by the QKG runtime for entity lookup and UMLS-
backed synonym matching.
This artifact is intended to be loaded into MongoDB collection:
primeKG.entities
Paper
This artifact is released with the paper:
Yao Wang, Zixu Geng, and Jun Yan. Quantum… See the full description on the dataset page: https://huggingface.co/datasets/HKAI-Sci/qkg-primekg-entities-with-cui.PrimeIntellect__INTELLECT-1-details
Dataset Card for Evaluation run of PrimeIntellect/INTELLECT-1
Dataset automatically created during the evaluation run of model PrimeIntellect/INTELLECT-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/PrimeIntellect__INTELLECT-1-details.PrimeVul_derived_finetuneFinetune dataset sample
sentinel-prime-350m-evals
Dataset Card for Evaluation run of qubitpage/sentinel-prime-350m
Dataset automatically created during the evaluation run of model qubitpage/sentinel-prime-350m
The dataset is composed of 7 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/qubitpage/sentinel-prime-350m-evals.PrimeVulNormForMultiVDprimevul_for_ml
