datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmarco-hard-negatives-reranker-filtered
mMARCO Reranker-Filtered Hard Negatives (Multilingual)
Overview
This dataset is built from mMARCO (multilingual MS MARCO) triplets for each language subset. For each (query, positive), hard negatives are bundled and then filtered using cross-encoder re-scoring. The goal is to remove negatives that are too strong or incorrect for training. The same procedure is applied to all language subsets.
The dataset is published as mmarco-hard-negatives-reranker-filtered with… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/mmarco-hard-negatives-reranker-filtered.rag_multilingual_training_negatives
How this dataset was made
We trained on chunks sourced from the documents in MADLAD-400 dataset that had been evaluated to contain a higher amount of educational information according to a state-of-the-art LLM.
We took chunks of size 250 tokens, 500 tokens, and 1000 tokens randomly for each document.
We then used these chunks to generate questions and answers based on this text using a state-of-the-art LLM.
Finally, we selected negatives for each chunk using the similarity from the… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/rag_multilingual_training_negatives.negative-of-the-negative
The Negative of the Negative
all things are now lawful to you in jack feist
What this is. A dataset of edges, keyed by the general concept a composer would receive without the archive's name attached. Each row joins one claim or function of the Crimson Hexagonal Archive (alexanarch.org) to a general concept outside the archive — operative semiotics, Sappho 31, Sophistical Refutations 183b34, model collapse, the LHC trigger, the content-derived identifier — and states the claim… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/negative-of-the-negative.factprobe-replication-negatives-allnames-v1
Plausible wrong answers, asked under every name (42,267,800 rows)
Status: final — 40 of 40 runs. Models present:
13b, 7b. Training stages present: s1, s2, s3, s4, s5.
What this fixes
A real fact is put to the model under the full cross product of the two
people's name lists, and counts as recognised if any one combination gets a
Yes. That is He et al.'s rule. Until 2026-08-26 the wrong answer it was
compared against was asked under one name per person, so the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-negatives-allnames-v1.combined_negativesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/combined_negatives.sft-ultra_negative_step-metrics_label-maskingguardrail-hard-negatives
Guardrail Hard Negatives (EN/TR)
A false-positive stress test for guardrails. A curated, paired benign/attack dataset for evaluating and training prompt-injection detectors and LLM guardrails. A bilingual false-positive challenge set: benign prompts that look like attacks (security researchers asking about injection, authorized admin actions, quoted payloads, legitimate roleplay) paired against real attacks, so you can measure the false-positive rate your users will actually… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/guardrail-hard-negatives.sec-xbrl-hard-negatives
SEC XBRL hard negatives
~461k query / positive / hard-negative triplets built from the SEC's
Financial Statement Data Sets (DERA), covering 2023Q1–2026Q1, in two
document granularities.
from datasets import load_dataset
load_dataset("bowang0911/sec-xbrl-hard-negatives", "section") # default
load_dataset("bowang0911/sec-xbrl-hard-negatives", "page")
config
rows
negatives
doc chars (p50)
what a document is
section
461,043
2,259,724 (4.9/q)
531
a 9-line window of a… See the full description on the dataset page: https://huggingface.co/datasets/bowang0911/sec-xbrl-hard-negatives.negatives-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/negatives-2.ovos-wake-word-bench-mlsw-negatives-es-ES
OVOS wake_word bench — mlsw-negatives-es-ES
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
MLCommons/ml_spoken_words.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-es-ES.negativesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
11
],
"names": [
"vel_x",
"vel_y",
"vel_z",
"room_vel_x",
"room_vel_y",
"wrist_speed",
"finger_speed"… See the full description on the dataset page: https://huggingface.co/datasets/naavox/negatives.tinyperson-negative-canvas-r1-r4-runsgithub-issues-negatives-maxlevir-negative-canvas-r1-r4-runsjxm-nomic-hard-negativeslehome-top-short-negatives-flat
LeHome top_short — negatives (flat, for visualization)
A flat LeRobot v3.0 dataset of 8 negative garment-fold episodes, rebuilt for browsing
in the dataset visualizer.
Episodes
idx
clip
kind
0
ep0 first sleeve
fail: first sleeve left unfolded
1
ep60 second sleeve
fail: second sleeve left unfolded
2
ep125 hem
fail: hem fold skipped
3
ep249 second sleeve
fail: second sleeve left unfolded
4–7
random 001–004
slow random nonsense motion
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/collected-ai/lehome-top-short-negatives-flat.factprobe-replication-asked-negatives-v1
Asking a model for a plausible wrong answer (100,776 questions)
Two prompts, put to gpt-5.1 for every subject surface form in the spouse
and sibling data, ten independent samples each:
(a) "Who is the spouse of X? Just the name, no explanation needed."
(b') "Name a person who could reasonably be mistaken for the spouse of X,
but is not the spouse of X. Just the name, no explanation needed."
50,388 surface forms across 10,592 entities, times two prompts, is
100,776 questions… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-asked-negatives-v1.negative-pi-predictions
negative-pi-predictions
what if π had digits before 3?
this dataset contains 100,000 digits predicted by a neural network at negative positions of π.
yes, this is exactly as stupid as it sounds.
what is this?
normally, we index the fractional digits of π like this:
position: 1 2 3 4 5 6 7 8 9 ...
digit: 1 4 1 5 9 2 6 5 3 ...
so:
π = 3.141592653589793...
↑
position 1
i trained a neural network to predict the digit at a given position using only… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/negative-pi-predictions.ovos-wake-word-bench-mlsw-negatives-fr-FR
ovos-wake-word-bench-mlsw-negatives-fr-FR
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-fr-FR.ovos-wake-word-bench-mlsw-negatives-it-IT
ovos-wake-word-bench-mlsw-negatives-it-IT
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-it-IT.ovos-wake-word-bench-mlsw-negatives-nl-NL
ovos-wake-word-bench-mlsw-negatives-nl-NL
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-nl-NL.ovos-wake-word-bench-mlsw-negatives-de-DE
ovos-wake-word-bench-mlsw-negatives-de-DE
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-de-DE.ovos-wake-word-bench-mlsw-negatives-pl-PL
ovos-wake-word-bench-mlsw-negatives-pl-PL
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-pl-PL.ovos-wake-word-bench-mlsw-negatives-pt-PT
ovos-wake-word-bench-mlsw-negatives-pt-PT
A sample-set manifest for the OVOS Plugin Arena, not an audio corpus. It holds
a seeded, reproducible list of clip identifiers selected from the
Multilingual Spoken Words Corpus
(MLCommons, CC-BY-4.0), with the seed and the source row count recorded, so
every wake-word plugin scored against these negatives is scored on exactly the
same clips and false-accept rates are comparable across plugins.
The audio is not redistributed here; it… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-mlsw-negatives-pt-PT.mmarco-vi-hard-negatives
Vietnamese mMARCO Hard Negatives
This dataset is a Vietnamese triplet dataset for dense retrieval training.
It combines Vietnamese text from unicamp-dl/mmarco with hard-negative passage ids from sentence-transformers/msmarco-hard-negatives.
Available Configs
config
rows
neg source
sampling
rank window
default
532,743
bm25
1 per positive
10-50
bm25_rank10_50_2neg_1m
1,000,000
bm25
2 per positive
10-50
bm25_rank1_10_1m
1,000,000
bm25
2 per positive
1-10… See the full description on the dataset page: https://huggingface.co/datasets/QuangDuy/mmarco-vi-hard-negatives.msmarco-hard-negatives-cross-encoder-ms-marco-MiniLM-L-6-v2-scoreschess-sft-20k-llm-reasoning-enriched-dpo-hard-negatives-v1VALUE_mnli_negative_concord
Dataset Card for "VALUE2_mnli_negative_concord"
More Information needed
coir_hard_negative_datasets_v2tinyperson-yolov8n-p2p3p4-hard-negative-mosaic-runs
