datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mia-fr-prompts
MIA-FR: French prompts for observing AI visibility
25 French prompts · 11 themes · 5 intent categories · Corpus v1.0 · CC BY 4.0
Lire en français · Field dictionary · Method and limitations · Source protocol
MIA-FR is a small, fixed prompt corpus published by Bertrand Morel / Edikka to support repeated observation of brand mentions and source citations in AI search interfaces. It gives practitioners a documented starting point for a French-language collection workflow.
This… See the full description on the dataset page: https://huggingface.co/datasets/edikka-lab/mia-fr-prompts.RUT-Bench
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions
This repository contains the RUT-Bench benchmark, which consists of 1638 test samples for evaluating LLM agents under realistic user interactions.
Paper: Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions
Code: GitHub
Collection: Hugging Face Collection
📖 Overview
RUT-Bench is a dedicated benchmark designed to assess… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/RUT-Bench.tb-explore17-mcode-m3-harness-variance
Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance
Three complete 17-task runs of the same dataset ref with the same agent and
model, differing only in execution substrate and concurrency, plus one isolated
rerun. The point of the bundle is not the resolve rate — it is how much the
resolve rate moves when nothing about the task or the model changes.
Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.shadow-llm-mia-signals
Shadow LLM MIA Signals (OLMo-2-1B)
Membership Inference Attack (MIA) signal tensors extracted from 128 shadow models
fine-tuned from allenai/OLMo-2-0425-1B.
Overview
This dataset enables research on membership inference attacks against large language models.
Each of 128 shadow models was trained on a different random subset of 64 out of 128 candidate
documents from the OLMo-mix-1124 pretraining dataset.
For each (model, document) pair, we extracted softmax prediction… See the full description on the dataset page: https://huggingface.co/datasets/matthewwicker/shadow-llm-mia-signals.swe-explore-find-dev100-runs
SWE-Explore find — dev-100 evidence runs
Five complete 100-trial runs of the find (fault-localization) arm of SWE-Explore,
kept because each one is load-bearing evidence for a specific claim about the
scoring fixes on branch fix/swe-explore-find-scoring of harness_bench.
Every run is 100 trials of the same dev-100 find subset, run through
Harbor with the swe_explore.agents:PiSut agent. index.jsonl
has one row per trial (500 rows); the full raw Harbor trial directories are in… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/swe-explore-find-dev100-runs.CheatClean
