emnlp
Datasets
All datasets matching “emnlp”Clotho-Moment
Clotho-Moment
This repository provides wav files used in Language-based Audio Moment Retrieval.
Each sample includes long audio containing some audio events with the temporal and textual annotation.
Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/
Code: https://github.com/line/lighthouse
Split
Train
train/train-{000..715}.tar
37930 audio samples
Valid
valid/valid-{000..108}.tar
5741 audio samples
Test
test/test-{000..142}.tar
7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.geodml-emnlp-2026
GEODML — EMNLP 2026 full reproducibility dataset
Companion data + code for the EMNLP 2026 submission
"Causal Analysis of LLM Search Rerankers via Double/Debiased Machine
Learning." Contains every input table, every fitted-model output,
every figure-generation script, and the raw LLM rerank outputs
needed to reproduce the paper end-to-end.
Condensed reviewer pack (5.6 MB, fast verification) →
ValerianFourel/geodml-emnlp-2026-reviewer
This dataset (1.8 GB, full pipeline) →… See the full description on the dataset page: https://huggingface.co/datasets/ValerianFourel/geodml-emnlp-2026.vibench
VIBench
VIBench is a benchmark for measuring vertical integration bias in
direct and agentic code generation. It contains 20 direct scenarios and
20 aligned agentic workflows covering realistic software integration
choices across cloud and API ecosystems. The bundle includes the
benchmark tasks, model metadata, system prompts, provider-proof
artifacts, and the blind detector-audit sample used for validation.
Viewer splits
The dataset viewer is intentionally simplified to… See the full description on the dataset page: https://huggingface.co/datasets/vibench-emnlp26/vibench.Clotho-Moment_CLAP_features
Clotho-Moment CLAP features
This repository contains audio and text features of Clotho-Moment dataset extracted by CLAP.
Using these features, we can reproduce the audio moments retrieval using Clotho-Moment, which is used in lighthouse.
Please also check demo page.
How to Download?
Run the following script:
from huggingface_hub import snapshot_download
repo_id = "lighthouse-emnlp2024/Clotho-Moment_CLAP_features"
local_dir = "./"
downloaded_path =… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment_CLAP_features.emnlp-2026-ifr-mair-full
EMNLP 2026 Findings: On-Policy Distillation Meets Off-Policy GRPO — Training Compact Instruction-Following Rerankers
Full MAIR set used for evaluation — 126 tasks and 9,356 queries (~8.8M candidate documents), spanning heterogeneous instruction-following retrieval domains.
Overview
This dataset is the complete MAIR (Massive Instructed Retrieval Benchmark) collection used as the out-of-distribution evaluation suite in the paper. It bundles all 126 MAIR tasks in a… See the full description on the dataset page: https://huggingface.co/datasets/anonymousauthor01/emnlp-2026-ifr-mair-full.vibench-results
VIBench Results
VIBench Results contains the retained raw generations,
detector-labeled outputs, complete runs, option-order and runtime
ablations, paper-facing summaries, figures, configs, audit files, and
static explorer indices used in the study.
The main paper evaluation covers 13 models, 15,600 direct generations,
and 2,000 agentic runs. Direct and agentic VIB are reported as
scenario-matched, share-normalized differences in affiliated-ecosystem
selection relative to strict… See the full description on the dataset page: https://huggingface.co/datasets/vibench-emnlp26/vibench-results.
