CoolFace
Datasetpublic

hrjinbb12345/search-swe-development

Search-SWE task development inputs Temporary public development inputs for one Search-SWE task submission. They are staged here so the submission manifest can pin an immutable revision while the task is under review. This is not a permanent official dataset path. Contents development/task-1-x-1/history.jsonl — the meeting transcripts the task uses development/task-1-x-1/validation/queries.jsonl — public development questions… See the full description on the dataset page: https://huggingface.co/datasets/hrjinbb12345/search-swe-development.

sourceHugging Facecc-by-4.0updated 5d agoView on Hugging Face
0likes38downloads
Dataset Card

Search-SWE task development inputs

Temporary public development inputs for one Search-SWE task submission. They are staged here so the submission manifest can pin an immutable revision while the task is under review. This is not a permanent official dataset path.

Contents

  • development/task-1-x-1/history.jsonl — the meeting transcripts the task uses
  • development/task-1-x-1/validation/queries.jsonl — public development questions
  • development/task-1-x-1/validation/golden_answers.jsonl — public reference answers and required factual points
  • development/task-1-x-1/validation/evidence.jsonl — transcript excerpts supporting the public answers

Source and licence

Derived from the ICSI Meeting Corpus (https://groups.inf.ed.ac.uk/ami/icsi/), released under the Creative Commons Attribution 4.0 licence (CC BY 4.0). The licence text is included here as ICSI_LICENSE.html.

Please cite: A. Janin, D. Baron, J. Edwards, D. Ellis, D. Gelbart, N. Morgan, B. Peskin, T. Pfau, E. Shriberg, A. Stolcke, C. Wooters, The ICSI Meeting Corpus, ICASSP 2003.

Provenance and processing

A cleaned subset of 67 meetings from the Bmr, Bro, and Bed series: 53,600 utterances, 12,155,329 bytes of UTF-8 dialogue text counting one newline per utterance. Speaker identifiers are normalised. Transcription artefacts such as interruption markers, fillers, and explicit uncertainty annotations are preserved.

The 30 public question-answer labels in validation/ were authored for this benchmark from the same transcripts and are not original ICSI annotations. A disjoint set of hidden questions and labels is used for evaluation and is not part of this dataset.