datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
miamMultilingual dIalogAct benchMark is a collection of resources for training, evaluating, and
analyzing natural language understanding systems specifically designed for spoken language. Datasets
are in English, French, German, Italian and Spanish. They cover a variety of domains including
spontaneous speech, scripted scenarios, and joint task completion. Some datasets additionally include
emotion and/or sentimant labels.ChineseNovels
中文小说数据集
包含内容:
网游/系统/重生
言情小说
同人/耽美小说
科幻小说
军事小说
以上加起来共4万本左右
海棠文学城小说:约1000本(未清洗)
mia-fr-prompts
MIA-FR: French prompts for observing AI visibility
25 French prompts · 11 themes · 5 intent categories · Corpus v1.0 · CC BY 4.0
Lire en français · Field dictionary · Method and limitations · Source protocol
MIA-FR is a small, fixed prompt corpus published by Bertrand Morel / Edikka to support repeated observation of brand mentions and source citations in AI search interfaces. It gives practitioners a documented starting point for a French-language collection workflow.
This… See the full description on the dataset page: https://huggingface.co/datasets/edikka-lab/mia-fr-prompts.mia-meeting
MIA Meeting E2E Dataset
Synthetic meeting dataset for end-to-end experiments:
audio to transcript
transcript plus roster to action items
action item extraction benchmark
Splits
train: 200 samples, 0 with linked audio
validation: 5 samples, 5 with linked audio
eval: 205 samples, 5 with linked audio
Structure
data/*.jsonl # split manifests
audio/<split>/* # linked audio files when available
transcripts/<split>/*.json #… See the full description on the dataset page: https://huggingface.co/datasets/minhthien/mia-meeting.RUT-Bench
Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions
This repository contains the RUT-Bench benchmark, which consists of 1638 test samples for evaluating LLM agents under realistic user interactions.
Paper: Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions
Code: GitHub
Collection: Hugging Face Collection
📖 Overview
RUT-Bench is a dedicated benchmark designed to assess… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/RUT-Bench.tb-explore17-mcode-m3-harness-variance
Terminal-Bench 2.1 explore-17 — mcode / MiniMax-M3 harness variance
Three complete 17-task runs of the same dataset ref with the same agent and
model, differing only in execution substrate and concurrency, plus one isolated
rerun. The point of the bundle is not the resolve rate — it is how much the
resolve rate moves when nothing about the task or the model changes.
Same everywhere: dataset ai-solution-finetune/terminal-bench-2-1-explore-17 at… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/tb-explore17-mcode-m3-harness-variance.shadow-llm-mia-signals
Shadow LLM MIA Signals (OLMo-2-1B)
Membership Inference Attack (MIA) signal tensors extracted from 128 shadow models
fine-tuned from allenai/OLMo-2-0425-1B.
Overview
This dataset enables research on membership inference attacks against large language models.
Each of 128 shadow models was trained on a different random subset of 64 out of 128 candidate
documents from the OLMo-mix-1124 pretraining dataset.
For each (model, document) pair, we extracted softmax prediction… See the full description on the dataset page: https://huggingface.co/datasets/matthewwicker/shadow-llm-mia-signals.SSAE-Dataset
Dataset Card
This is the official dataset repository for the paper "Step-Level Sparse Autoencoder for Reasoning Process Interpretation".
Paper: Arxiv
Code: GitHub
Collection: HuggingFace
Dataset Overview
The repository hosts three distinct datasets covering the domains of mathematical reasoning and code generation. Each subset is pre-partitioned into training and validation splits to facilitate reproducible experiments.
1. GSM8K (Math)
Description: A… See the full description on the dataset page: https://huggingface.co/datasets/Miaow-Lab/SSAE-Dataset.swe-explore-find-dev100-runs
SWE-Explore find — dev-100 evidence runs
Five complete 100-trial runs of the find (fault-localization) arm of SWE-Explore,
kept because each one is load-bearing evidence for a specific claim about the
scoring fixes on branch fix/swe-explore-find-scoring of harness_bench.
Every run is 100 trials of the same dev-100 find subset, run through
Harbor with the swe_explore.agents:PiSut agent. index.jsonl
has one row per trial (500 rows); the full raw Harbor trial directories are in… See the full description on the dataset page: https://huggingface.co/datasets/miaomiao64/swe-explore-find-dev100-runs.task896_miam_language_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task896_miam_language_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task896_miam_language_classification.CheatCleannutriswap-healthy-food-alternatives
