datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ruler-100-nemotron
RULER-100 — Nemotron-Nano-v3 tokenized
RULER long-context evaluation data, regenerated with the
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (instruct) tokenizer so the labeled context
lengths are exact for that model — instead of drifting, as they do when RULER data tokenized for a
different model (e.g. Qwen3) is fed to Nemotron.
What's here
7 context lengths: 4096, 8192, 16384, 32768, 65536, 131072, 262144 (the model's max).
13 RULER tasks: niah_single_1/2/3… See the full description on the dataset page: https://huggingface.co/datasets/jet-ai/ruler-100-nemotron.STXBP1-RAG-Nemotron
🧬⚡ STXBP1-ARIA RAG Database v10.1 - NVIDIA Nemotron Embeddings
The most advanced RAG database for STXBP1 therapeutic research.
A pre-built ChromaDB vector database containing:
571,816 indexed text chunks from ~17,000 curated PubMed Central (PMC) biomedical papers + 165 base editing analysis entries,
(https://huggingface.co/datasets/SkyWhal3/stxbp1-base-editing-sweep),
embedded with NVIDIA's state-of-the-art Llama-Nemotron-Embed-1B-v2 model featuring 2048-dimensional embeddings.
⚡… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/STXBP1-RAG-Nemotron.
