hrjinbb12345/search-swe-development
Search-SWE task development inputs Temporary public development inputs for one Search-SWE task submission. They are staged here so the submission manifest can pin an immutable revision while the task is under review. This is not a permanent official dataset path. Contents development/task-1-x-1/history.jsonl — the meeting transcripts the task uses development/task-1-x-1/validation/queries.jsonl — public development questions… See the full description on the dataset page: https://huggingface.co/datasets/hrjinbb12345/search-swe-development.
Search-SWE task development inputs
Temporary public development inputs for one Search-SWE task submission. They are staged here so the submission manifest can pin an immutable revision while the task is under review. This is not a permanent official dataset path.
Contents
development/task-1-x-1/history.jsonl— the meeting transcripts the task usesdevelopment/task-1-x-1/validation/queries.jsonl— public development questionsdevelopment/task-1-x-1/validation/golden_answers.jsonl— public reference answers and required factual pointsdevelopment/task-1-x-1/validation/evidence.jsonl— transcript excerpts supporting the public answers
Source and licence
Derived from the ICSI Meeting Corpus (https://groups.inf.ed.ac.uk/ami/icsi/), released under the Creative Commons Attribution 4.0 licence (CC BY 4.0). The licence text is included here as ICSI_LICENSE.html.
Please cite: A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, C. Wooters, The ICSI Meeting Corpus, ICASSP 2003.
Provenance and processing
A cleaned subset of 67 meetings from the Bmr, Bro, and Bed series: 53,600 utterances, 12,155,329 bytes of UTF-8 dialogue text counting one newline per utterance. Speaker identifiers are normalised. Transcription artefacts such as interruption markers, fillers, and explicit uncertainty annotations are preserved.
The 30 public question-answer labels in validation/ were authored for this benchmark from the same transcripts and are not original ICSI annotations. A disjoint set of hidden questions and labels is used for evaluation and is not part of this dataset.
