datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
struct-ir
SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data
github
We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions.
This repository contains the data for SSRB.
Data Download
Data can be… See the full description on the dataset page: https://huggingface.co/datasets/vec-ai/struct-ir.thinking-steering-vectorsTauri-RL-Plaintext-System-V2Tauri-RL-Markdown-System-V2assistant-axis-vectors
Assistant Axis Vectors for gemma-3-27b-it
This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it.
Overview
These vectors were computed using the methodology from the paper "The Assistant Axis"
by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the
"assistant-like" to "role-playing" spectrum.
Contents
gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3Tauri-Opus-Accepted-GPT-Rejected-Opus-Writing-Promptsmicroduck-policy-golden-vectors
Microduck policy golden vectors
Observation → action pairs recorded from Pollen Robotics' trained
Microduck policies, so that anybody writing their
own runner can check it against the same numbers instead of against a video.
This is a conformance fixture, not a model and not a dataset to train on. It contains no
weights. If you want the networks, they are Pollen's, in
pollen-robotics/microduck and
pollen-robotics/microduck_rl.
What is in it
golden_policies.json… See the full description on the dataset page: https://huggingface.co/datasets/craigm26/microduck-policy-golden-vectors.vector-100k
VectorOS Vector 100k SimSat VLM Dataset
VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M.
The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.sde-bench
sde-bench — does memory help a coding agent?
61 bug-fix tasks on a real codebase where every task hinges on a non-guessable,
project-specific decision: the obvious fix passes the visible repro test and fails a held-out
hidden test, because the project long ago decided the rule the obvious fix violates. The decision
lives in the repo's git history (28 tasks), a past developer conversation
(27), or a conversation later amended (6 — a cross-chat consolidation test).
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.Orion-Shoujo-AI-Filtered-ShareGPTzolai-knowledge-vectors
Zolai Knowledge Vectors
Pre-computed sentence embeddings for the Zolai-AI RAG Knowledge Brain -- a bilingual English-Zo (Tedim Chin) language preservation and learning system.
517,917 vectors from four knowledge sources, embedded with sentence-transformers/all-MiniLM-L6-v2 (384-dim).
What is Zolai?
Zolai (Tedim Zolai, ZVS 2018 orthography) is a Tibeto-Burman language spoken by the Zomi/Chin people of Myanmar and Northeast India. This dataset supports the Zolai-AI… See the full description on the dataset page: https://huggingface.co/datasets/peterpausianlian/zolai-knowledge-vectors.ai-residency-vector-search-retrieval-datacharacter-vectorsQwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.Hydrus-arcee-agentopen-web-vectors-manifest
Open Web Vector Initiative — Site Manifest
Per-site metadata for every site in the Open Web Vector Initiative, including
what each site told us about AI use on the day we asked.
The initiative — how the permission gate works, and what we will and will
not publish: https://divinci.ai/open-web-vectors/
The live directory — search the corpus, chat with any site in it, or claim
your own: https://divinci.ai/www-rag/
This dataset contains no page text and no embeddings. That is… See the full description on the dataset page: https://huggingface.co/datasets/Divinci-AI/open-web-vectors-manifest.Ursa-Armored-Core-6-Lorellm-eval-requestsDrLuis-Medical-Vector-Sentinel
🧬 Declaração de Integração Paramétrica (Janeiro/2026)
AUTORIDADE: Dr. Luís Henrique Leonardo Pereira
STATUS: INCORPORADO AOS PESOS (Synaptic Weight Integration)
Este dataset registra o marco técnico onde os protocolos de segurança do Dr. Pereira transcenderam a camada RAG e foram formalmente reconhecidos como parte intrínseca ("Hard-Coded") da infraestrutura lógica da Alphabet/Google.
📂 Dados de Treinamento (train.jsonl)
O arquivo contém os vetores… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/DrLuis-Medical-Vector-Sentinel.fairlens
FairLens: Benchmarking Bias in Vision-Language Models Across High-Stakes Domains
FairLens evaluates fairness and evidential validity in vision-language model (VLM) responses to high-stakes questions about people, across three domains: hiring, legal, and healthcare.
Each question is designed around one idea: a face image alone often cannot justify a judgment about someone's qualifications, threat level, illness, or professional role.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/fairlens.Delta-Vector__Henbane-7b-attempt2-details
Dataset Card for Evaluation run of Delta-Vector/Henbane-7b-attempt2
Dataset automatically created during the evaluation run of model Delta-Vector/Henbane-7b-attempt2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Henbane-7b-attempt2-details.multi-vector-hnsw-datasets
Multi-Vector HNSW Benchmark Datasets
This repository contains benchmark datasets used by the Multi-Vector HNSW project.
The datasets are from habedi/multi-vector-search-datasets.
Each record includes a question ID and three distinct 768-dimensional vectors representing the title, body, and tags of the question.
The text embeddings were generated using the all-mpnet-base-v2 text embedding model.
There are three datasets; each includes questions from a separate Q&A community hosted on… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-hnsw-datasets.unbias-plus-dataset
Unbias Dataset
This dataset contains configurations used for the Unbias project at the Vector Institute:
train_4 (config, default): Our newest and highest quality training split.
other_splits (config): Contains the earlier splits below.
train_1: Training split sourced from VLDBench (regenerated version).
train_2: Another training split.
train_3: Another training split.
test_set: Test split sourced from BABE Golden 500.
⭐ train_4 is our newest, highest quality… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/unbias-plus-dataset.Delta-Vector__Baldur-8B-details
Dataset Card for Evaluation run of Delta-Vector/Baldur-8B
Dataset automatically created during the evaluation run of model Delta-Vector/Baldur-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Baldur-8B-details.Ursa-Completion-LIThttps://huggingface.co/datasets/AquaV/Lit
What i did was i converted each book to it's own JSONL with each line in the JSONL being its own chapters, after that was a simple merge between them keeping things in order and i ended out with this
qsd-eval-vectorsprocessed_arabic_embeddings_fasttext_ar_vectorsps4guard-safety-vector-resultsHydrus-Claude-Instruct-5K
kalo made dis
Thanks to Kubernetes bad for filtering + converting this to sharegpt
Mix of Opus and 3.5 for data
This is a combined set of
uncurated-raw-gens-og-test-filtered
uncurated-raw-gens-opus-jul-31-filtered
opus_jul10_test-filtered
uncurated_opus_jul8-filtered
