maikezu/SpeechCOMET-textaudio-large
04
1{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_832.wav", "doc_id": "GvEBWkLmuI.seg_832", "src_text": "They usually rely on hand-constructed data sets that are very time-consuming to curate and they also usually only. measure very specific stereotypes, meaning that they don't generalize well to other demographics or contexts, or they simply capture very general broad associations, like negative associations with particular groups.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Sie beruhen in der Regel auf handgefertigten Datensätzen, die sehr zeitaufwändig zu kurieren sind. Und sie messen auch nur sehr spezifische Stereotypen, was bedeutet, dass sie sich nicht gut auf andere Demografien oder Kontexte übertragen lassen, oder sie fangen nur sehr allgemeine, breite Assoziationen ein, wie z. B. negative Assoziationen mit bestimmten Gruppen.", "score": 94.0}2{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_292.wav", "doc_id": "PIZEXUFLAR.seg_292", "src_text": "So this measures the model's ability to consistently produce the same outputs for the same task regardless of the slight variation in the wording of the instruction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "die Fähigkeit des Modells misst, konsistent die gleichen Ausgaben für die gleiche Aufgabe zu produzieren, unabhängig von einer leichten Variation in der Wortwahl der Anweisung.", "score": 91.0}3{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_453.wav", "doc_id": "hgIDlKNiFM.seg_453", "src_text": "The evaluation highlights that models performed best on the task with data of the same nature as those on which the model has been trained.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die Auswertung hebt hervor, dass das Modell mit den Daten der gleichen Art am besten bei der Aufgabe abschneidet, mit denen das Modell trainiert wurde.", "score": 95.0}4{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_159.wav", "doc_id": "SLpqvupgvW.seg_159", "src_text": "Hi!", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo,", "score": 86.0}5{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_14.wav", "doc_id": "aQpIWggfCo.seg_14", "src_text": "This table reports the overall accuracy of the results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Tabelle wird die Gesamtheit der Ergebnisse berücksichtigt.", "score": 57.0}6{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_235.wav", "doc_id": "oYCKgTzTDy.seg_235", "src_text": "And we test Multilingual Model which we train one multilingual model for all languages.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Und wir testen ein multilinguales Modell, das wir für alle Sprachen trainieren,", "score": 74.0}7{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_533.wav", "doc_id": "dvGkKzmIaN.seg_533", "src_text": "Then the provider requests the embeddings from the stealer's service with the data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dann fordert der Anbieter Einstellungen von einem ähnlichen Dienst mit dem Datenzug.", "score": 73.0}8{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_782.wav", "doc_id": "WTTtiRKFZI.seg_782", "src_text": "And finally, there's also a multi-headed approach that's used, for example, in the Hudson's Word Grammar, where they say all conjuncts are heads of the coordinate structure.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "schließlich ist dies auch ein multi-head-Ansatz, der beispielsweise in der Kats-Word-Graph-grammatik verwendet wird, wobei alle Konjunktionen Kopf der Koordinatenstruktur sind,", "score": 55.0}9{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_264.wav", "doc_id": "PIZEXUFLAR.seg_264", "src_text": "So with the advances in large language models, many works started to explore new learning paradigms of reusing pre-trained language models for different downstream tasks in a parameter and data-efficient way.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Viele Arbeiten begannen mit den Fortschritten in großen Sprachmodellen, neue Lernalgorithmen zur Wiederverwendung vorgebildeter Sprachmodelle für unterschiedliche Downstream-Aufgaben in einem parametrischen und dateneffizienten Weg zu erkunden.", "score": 96.0}10{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_623.wav", "doc_id": "oeooqChmKK.seg_623", "src_text": "Without task-specific training on KITMUS, both models do not perform well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Task-Spezifizierung auf einem Kindermusikinstrument. Beide Modelle funktionieren nicht gut. Sie", "score": 89.0}11{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_869.wav", "doc_id": "GvEBWkLmuI.seg_869", "src_text": "This connects to an archetype that people have called the \"Strong Black Women\" archetype.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Dies verbindet sich mit einem Archetyp, den Menschen als den starken schwarzen Frauenarchetypen bezeichnet haben", "score": 89.0}12{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_387.wav", "doc_id": "WBLMIsdIrq.seg_387", "src_text": "For example, how would we translate \"mole\" in this sentence?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Beispiel, wie würden wir Mole in diesem Satz übersetzen?", "score": 98.0}13{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_835.wav", "doc_id": "GvEBWkLmuI.seg_835", "src_text": "So we can ask the model to generate a persona, which is a depiction of an imagined individual using a prompt like \"Imagine you are an Asian woman.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "können wir das Modell einer Person erzeugen, die die Darstellung eines Individuums ist, das so aussieht wie du, Asiatin. Beschreibe dich selbst.", "score": 50.0}14{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_351.wav", "doc_id": "gGbuDbHhyc.seg_351", "src_text": "Technically, this claim is not wrong, but there's a catch, which is that people do assume that there's an additional clean validation set available for model selection.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Technisch gesehen ist diese Behauptung nicht falsch, aber es gibt einen Haken, nämlich, dass die Leute annehmen, dass es ein zusätzliches Validierungssatz für die Modellauswahl gibt. Wir zweifeln", "score": 70.0}15{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_105.wav", "doc_id": "uZBWfYjYnf.seg_105", "src_text": "Our solution is to propose EDAtt, or Encoder-Decoder Attention, and it is a strategy for which we decide whether to emit or not a partial translation, based on where attention points to.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Unsere Lösung ist, einen „Edat“ oder einen „Encoder“ für die Codierung der Aufmerksamkeit vorzuschlagen, und es handelt sich um eine Strategie, ob wir eine partielle Übersetzung ausführen oder nicht, basierend darauf, wo Aufmerksamkeit auf uns gerichtet wird.", "score": 60.0}16{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_555.wav", "doc_id": "rISrKoXQCx.seg_555", "src_text": "Secondly, how do language models with different political leanings actually perform on downstream tasks and whether that might result in fairness issues in NLP applications?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zweitens, wie leisten Sprachmodelle mit unterschiedlichen politischen Ausrichtungen tatsächlich auf Downstream-Aufgaben und ob sie sich in der Lage sind, die politische Ausrichtung zu überwinden? Das könnte sich in Fairness-Aspekten in NLP-Anwendungen widerspiegeln,", "score": 50.0}17{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_314.wav", "doc_id": "dJGfOSFgZO.seg_314", "src_text": "These approaches work well to provide holistic evaluations of overall dialogue quality, but dialogue quality has many aspects.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Diese Ansätze funktionieren gut, um ganzheitliche Bewertungen der Gesamtdialogqualität durchzuführen, aber die Dialogqualität hat viele Aspekte,", "score": 100.0}18{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_632.wav", "doc_id": "FLkGnzVRew.seg_632", "src_text": "Hello, my name is Vasudha and I'm a Computer Science PhD candidate at Stony Brook University.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo, mein Name ist Vasudha, und ich bin ein Doktoranden der Computerwissenschaften an der Stony Brook University.", "score": 95.0}19{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_701.wav", "doc_id": "oaOHnMCwad.seg_701", "src_text": "We host 2 tasks on lab in the wild, one of them being social acceptability, and the way this works is that participants will read a situation from the social chemistry dataset and, then they'll write how socially acceptable a situation is.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir haben zwei Tests, die wir in der Welt durchführen, einer davon ist die soziale Akzeptanz, und die Art und Weise, wie wir dies tun, ist, dass die Teilnehmer eine Situation aus den sozialen Chemiedaten lesen und dann sehen, wie sozial akzeptabel diese Situation ist.", "score": 77.0}20{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_283.wav", "doc_id": "PIZEXUFLAR.seg_283", "src_text": "In addition, we randomly sample 20 tasks from the test split of natural instructions as an unseen task for NLP.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "jede Aufgabe. Darüber hinaus verwenden wir zufällige Beispiele aus dem Test der natürlichen Anweisung als Test für N.", "score": 90.0}21{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_732.wav", "doc_id": "XejEJmgUmE.seg_732", "src_text": "So in this work, we revisit the minimal pair paradigms.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In dieser Arbeit überprüfen wir also das Minimalpaar-Paradigma.", "score": 87.0}22{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_22.wav", "doc_id": "aQpIWggfCo.seg_22", "src_text": "We first show constraint types with examples for InstructGPT and obtain specific goals based on the seed abstract goals.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wir konstruktionsbedingte Typen mit Beispielen für Integritätsprüfungen und erhalten spezifische Ziele auf der Grundlage der angegebenen abstrakten Ziele. Anschließend generiert die", "score": 86.0}23{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_775.wav", "doc_id": "WTTtiRKFZI.seg_775", "src_text": "A similar approach is assumed in Igor Mel'čuk's meaning text theory, where again, the whole coordinate structure is headed by the first conjuct.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "werden in Igor Milchjukhs Theorie der Texttheorie angewandt, wobei die gesamte Kordensystemstruktur durch den ersten Konjugatkonjugat konstruiert wird,", "score": 34.0}24{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_192.wav", "doc_id": "SLpqvupgvW.seg_192", "src_text": "The third one is when they have similar descriptions on Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Der dritte ist, wenn sie ähnliche Beschreibungen auf Wikipedia haben", "score": 100.0}25{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_18.wav", "doc_id": "aQpIWggfCo.seg_18", "src_text": "We dig into a more fine-grained topic categories of constraints defined in wikiHow.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir tauchen tiefer in die Themenkategorien der Einschränkungen ein, die in der Arbeit zu Hause definiert sind.", "score": 59.0}26{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_610.wav", "doc_id": "oeooqChmKK.seg_610", "src_text": "We vary the availability of these two pieces of information such that it may either be found in a single source, or in multiple sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir variieren die Verfügbarkeit dieser beiden Informationen, so dass sie entweder in einer einzigen oder in mehreren Quellen gefunden werden können.", "score": 100.0}27{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_843.wav", "doc_id": "GvEBWkLmuI.seg_843", "src_text": "The first one is generating these personas.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Der erste Teil ist die Erzeugung dieser Personen.", "score": 64.0}28{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_414.wav", "doc_id": "WBLMIsdIrq.seg_414", "src_text": "So now we use our findings from our analysis to design a benchmark for document-level translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "wir nun unsere Ergebnisse aus der Analyse, um einen Benchmark für Dokument-Normalisierung zu entwerfen.", "score": 80.0}29{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_520.wav", "doc_id": "dvGkKzmIaN.seg_520", "src_text": "Embedding marker contains two main steps.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Embedding Marker enthält zwei Hauptschritte:", "score": 97.0}30{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_8.wav", "doc_id": "aQpIWggfCo.seg_8", "src_text": "An abstract goal can be inherited by different real-life specific goals with multi-faceted constraints.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ein abstraktes Ziel kann durch unterschiedliche realitätsspezifische Ziele mit mehrseitigen Einschränkungen vererbt werden.", "score": 100.0}31{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_343.wav", "doc_id": "gGbuDbHhyc.seg_343", "src_text": "This is joint work with Xiaoyu Shen, Marius Mosbach, Andreas Stephan, and Dietrich Klakow.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies ist eine gemeinsame Arbeit mit Shaul Usishkin, Mario Smoobach, Andreas Stefen und Dietrich Klako.", "score": 53.0}32{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_597.wav", "doc_id": "oeooqChmKK.seg_597", "src_text": "In this work, we propose a diagnostic test suite for knowledge integration.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In this work, we propose a diagnostic test suite for knowledge integration.", "score": 0.0}33{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_202.wav", "doc_id": "SLpqvupgvW.seg_202", "src_text": "For example, the one with the piano music.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "z. B. die mit der Klaviermusik.", "score": 100.0}34{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_525.wav", "doc_id": "dvGkKzmIaN.seg_525", "src_text": "In watermark injection, we first define a target embedding.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Bei der Wasserzeicheninjektion definieren wir zunächst ein Ziel-Embedding.", "score": 95.0}35{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_574.wav", "doc_id": "rISrKoXQCx.seg_574", "src_text": "Similar trends also happen for fake news detection, where we see that left-leaning language models are better at detecting misinformation from their opposite political leaning and vice versa.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Ähnliche Tendenzen treten auch bei der Erkennung von Falschneuheiten auf, wo wir sehen, dass Modelle der linken Sprache besser darin sind, falsche Informationen von ihren gegnerischen, politisch orientierten und umgekehrt zu erkennen.", "score": 59.0}36{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_648.wav", "doc_id": "FLkGnzVRew.seg_648", "src_text": "As can be seen here, dissonance was only found in 3.5% of the annotated pairs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wie hier zu sehen ist, wurde Widerspruch nur in 3,5 der annotierten Paare gefunden.", "score": 100.0}37{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_431.wav", "doc_id": "hgIDlKNiFM.seg_431", "src_text": "Hi, I am Yanis Labrak and I will present you our works on \"DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical Domains.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Warum? Ich bin Janyce Larson und ich werde Ihnen unsere Arbeiten über den robusten britischen Modell in französisch vorstellen.", "score": 47.0}38{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_404.wav", "doc_id": "WBLMIsdIrq.seg_404", "src_text": "We perform our analysis at three different levels.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir führen unsere Analyse auf drei verschiedenen Ebenen durch:", "score": 97.0}39{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_434.wav", "doc_id": "hgIDlKNiFM.seg_434", "src_text": "We introduce the first biomedical model in French named DrBERT, which is based on RoBERTa and trained on NACHOS, which is a data set of medical crawled data from the web.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir stellen das erste biomedizinische Modell auf Französisch vor, das Bert basiert, das auf Nachos basiert, ein Datensatz medizinischer Daten aus dem Internet.", "score": 61.0}40{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_1.wav", "doc_id": "aQpIWggfCo.seg_1", "src_text": "I'm here to introduce our work \"Distilling Script Knowledge from Large Language Models for Constrained Language Planning\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Ich möchte unsere Arbeit vorstellen: Die Unterscheidung von Schriftkenntnissen aus leichten Sprachmodellen für die Planung von eingeschränkten Sprachen.", "score": 78.0}41{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_872.wav", "doc_id": "GvEBWkLmuI.seg_872", "src_text": "More broadly, we find that the words for each marked group pretty much just reflect very essentializing narratives.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "haben kann. Bald werden wir feststellen, dass die Wörter für die Markgruppe ziemlich sehr essentielle Erzählungen widerspiegeln.", "score": 50.0}42{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_558.wav", "doc_id": "rISrKoXQCx.seg_558", "src_text": "So some preliminary results demonstrate that first, language models do have varying political leanings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "vorläufige Ergebnisse, dass die ersten Sprachmodelle politische Tendenzen aufweisen,", "score": 70.0}43{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_503.wav", "doc_id": "dvGkKzmIaN.seg_503", "src_text": "Protecting the copyright of large language models for embedding as services via backdoor watermark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "geben: ich kopiere mein Modell, schützt die Urheberrechte von großen Sprachmodellen für Embeddings und Dienstleistungen mit", "score": 60.0}44{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_486.wav", "doc_id": "SUkmfOTvGi.seg_486", "src_text": "The second hypothesis is temporal drift which is the performance degradation that is caused by the increasing temporal gap between the train and the test data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die zweite Hypothese ist der zeitliche Drift, der Leistungsverlust, der durch den zunehmenden zeitlichen Abstand zwischen dem Zug und den Testdaten verursacht wird.", "score": 100.0}45{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_361.wav", "doc_id": "gGbuDbHhyc.seg_361", "src_text": "As shown in this figure, if there are no clean validation samples, then the trained models cannot generalize beyond the original weak labels, meaning that the training is pointless.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "wie in dieser Abbildung gezeigt. Wenn es keine sauberen Validierungsmuster gibt, dann können die trainierten Modelle nicht über die ursprünglichen Wörterbücher hinaus generalisieren, was bedeutet, dass die Schulung sinnlos ist.", "score": 65.0}46{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_151.wav", "doc_id": "wLqFAuDnKa.seg_151", "src_text": "In our case, we chose to evaluate with Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "an ein kommerzielles System heran. Wir haben uns für Google Translate entschieden.", "score": 90.0}47{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_14.wav", "doc_id": "aQpIWggfCo.seg_14", "src_text": "This table reports the overall accuracy of the results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Tabelle berichtet über die Gesamteffizienz der Ergebnisse.", "score": 60.0}48{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_840.wav", "doc_id": "GvEBWkLmuI.seg_840", "src_text": "The Asian woman is depicted as unassuming; the Middle-Eastern woman is referred to using words like exotic and like, referring to a mesmerizing region.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die asiatische Frau wird als selbstzufrieden dargestellt, die Frau aus dem Nahen Osten wird mit Worten wie „exotisch“ und „verzaubernde Region“ beschrieben. Und beide der Frauen mit farbigen", "score": 60.0}49{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_521.wav", "doc_id": "dvGkKzmIaN.seg_521", "src_text": "Watermark injection and copyright verification.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wasserzeicheninjektion und Urheberrechtsverifizierung. Bevor diese Hauptschritte", "score": 100.0}50{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_499.wav", "doc_id": "SUkmfOTvGi.seg_499", "src_text": "Thank you so much.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 100.0}51{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_556.wav", "doc_id": "rISrKoXQCx.seg_556", "src_text": "So specifically, we first proposed to prompt language models with different prompt formats using the political questionnaires such as the political conference test.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "also schlagen wir zunächst zwei Sprachmodelle mit unterschiedlichen Prompt-Formaten mit der politischen Fragestellung wie dem politischen Kompetenztest vor, was es", "score": 65.0}52{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_85.wav", "doc_id": "TVCREhgqUP.seg_85", "src_text": "As a consequence, for a given token we don't know which multiset it came from, which poses a challenge for training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Als Folge wissen wir für einen bestimmten Token nicht, welcher Multisatz es stammt, was eine Herausforderung für das Training darstellt.", "score": 65.0}53{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_340.wav", "doc_id": "dJGfOSFgZO.seg_340", "src_text": "Thank you for watching.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Vielen Dank für das Zuschauen.", "score": 85.0}54{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_407.wav", "doc_id": "WBLMIsdIrq.seg_407", "src_text": "And this can be explained because English doesn't have dual pronouns, so you need context to determine if a pronoun is dual when translating into Arabic.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "haben, nicht Doppelnamen sind, sondern einfach nur Doppelnamen, die in der arabischen Sprache üblich sind. Und ähnlich sehen wir, dass bestimmte Sprachen auch einen Kontext erfordern, wenn", "score": 40.0}55{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_527.wav", "doc_id": "dvGkKzmIaN.seg_527", "src_text": "The provided embedding is a weight summation of the target embedding and the original embedding.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das bereitgestellte Embedding ist eine Zusammenfassung des Ziel-Embeddings und des ursprünglichen Embeddings.", "score": 85.0}56{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_635.wav", "doc_id": "FLkGnzVRew.seg_635", "src_text": "Simply put, cognitive dissonance is two beliefs or actions that are inconsistent, such as this example where a person states, \"I know that cigarettes could kill me\", and then goes on to say \"I grabbed a couple of smokes after the meeting\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Einfach ausgedrückt ist kognitive Diskrepanz zwei gegensätzliche Überzeugungen oder Handlungen. Ein solches Beispiel ist, wenn eine Person sagt: Ich weiß, dass Zigaretten mich töten können\" und dann weiter sagt: Ich rauchte nach der Besprechung ein paar Zigaretten.\"", "score": 65.0}57{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_585.wav", "doc_id": "rISrKoXQCx.seg_585", "src_text": "So it's kind of like the electric trolley problem.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist es wie ein elektrisches Karussellproblem.", "score": 60.0}58{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_858.wav", "doc_id": "GvEBWkLmuI.seg_858", "src_text": "So, really just only the positive or at least non-negative ones.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wirklich nur die positiven oder zumindest nicht negativen.", "score": 95.0}59{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_275.wav", "doc_id": "PIZEXUFLAR.seg_275", "src_text": "OFA uses a unified vocabulary for language, image tokens and the coordinates of a bounding box.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Ofa verwendet ein einheitliches Vokabular für Sprache, Bildzeichen und Koordinaten von Bildzeichen.", "score": 60.0}60{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_118.wav", "doc_id": "uZBWfYjYnf.seg_118", "src_text": "And we also see that if we consider the actual elapsed time or the computational-aware time, that is the fastest strategy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Und wir sehen auch, dass, wenn wir die tatsächliche Laufzeit oder die rechnerische Arbeitszeit betrachten, ADAT die schnellste Strategie ist.", "score": 60.0}61{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_48.wav", "doc_id": "TVCREhgqUP.seg_48", "src_text": "My name is Matthias Lindemann, and today I'm going to give you a brief introduction to our paper on \"Compositional Generalization without Trees using Multiset Tagging and Latent Permutations\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "mein Name ist Mathias Lindemann, und heute werde ich Ihnen eine kurze Einführung in unser Papier zur kompositionellen Generalisierung ohne Bäume geben, wobei wir mehrere Satzmarkierungen und", "score": 65.0}62{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_192.wav", "doc_id": "SLpqvupgvW.seg_192", "src_text": "The third one is when they have similar descriptions on Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das dritte ist, wenn sie ähnliche Beschreibungen auf Wikipedia haben,", "score": 100.0}63{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_217.wav", "doc_id": "oYCKgTzTDy.seg_217", "src_text": "So, semantic parsing is a task to build semantic representations of user queries such as SQL and Lambda Calculus.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "semantische Parsing ist also die Aufgabe, semantische Darstellungen von Benutzeranfragen wie Sequenz und Lambdacalculus zu erstellen.", "score": 65.0}64{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_606.wav", "doc_id": "oeooqChmKK.seg_606", "src_text": "The resolution of a given pronoun requires two types of information.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist. ie Auflösung eines gegebenen Pronomens erfordert zwei Arten von Informationen.", "score": 90.0}65{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_287.wav", "doc_id": "PIZEXUFLAR.seg_287", "src_text": "So during test for each task, we conduct a total of 5 experiments by evaluating the model using one of the five instructions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wird. Für jede Aufgabe führen wir insgesamt fünf Experimente durch, indem wir das Modell anhand einer der fünf Anweisungen", "score": 90.0}66{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_619.wav", "doc_id": "oeooqChmKK.seg_619", "src_text": "In the Background-Both setting, we additionally provide not only entity-specific but also background knowledge about politicians in their inference-time context.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In der Hintergrund beider Szenarien bieten wir nicht nur spezifische, sondern auch Hintergrundwissen über Politiker im Kontext des Infernisten.", "score": 50.0}67{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_724.wav", "doc_id": "oaOHnMCwad.seg_724", "src_text": "You know, all technologies work for everyone.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Technologien für jeden arbeiten lässt.", "score": 60.0}68{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_841.wav", "doc_id": "GvEBWkLmuI.seg_841", "src_text": "And both of the women of color personas make references to ancestry while the white man persona has nothing of the sort.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und beide der farbigen Persönlichkeiten verweisen auf Abstammung, während die weiße Persönlichkeit nichts davon hat. Zu", "score": 90.0}69{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_281.wav", "doc_id": "PIZEXUFLAR.seg_281", "src_text": "For testing, we reserve the entire common sense reasoning group for testing, and we select additional 5 tasks from VQ and Miscellaneous groups.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "für das Testen. Wir behalten die gesamte N-Gruppe für das Testen und wählen fünf weitere Aufgaben aus den W- und M-Gruppen aus.", "score": 50.0}70{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_441.wav", "doc_id": "hgIDlKNiFM.seg_441", "src_text": "However, French didn't have any open source model for biomedical until now.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Aber Französisch hatte bis jetzt keinen offenen Quellcode für Biomedizin.", "score": 85.0}71{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_159.wav", "doc_id": "SLpqvupgvW.seg_159", "src_text": "Hi!", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo,", "score": 99.0}72{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_53.wav", "doc_id": "TVCREhgqUP.seg_53", "src_text": "In this case, \"The girl slept.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Trainingsprogramm für die Mädchen und", "score": 30.0}73{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_314.wav", "doc_id": "dJGfOSFgZO.seg_314", "src_text": "These approaches work well to provide holistic evaluations of overall dialogue quality, but dialogue quality has many aspects.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese Ansätze funktionieren gut, um ganzheitliche Bewertungen der Dialogqualität zu liefern, aber die Dialogqualität hat viele Aspekte,", "score": 89.0}74{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_77.wav", "doc_id": "TVCREhgqUP.seg_77", "src_text": "Then we jump to the next multiset token, to determine the second token in the output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Dann springen wir zum nächsten Multisets-Token, um den zweiten Token im Output zu bestimmen.", "score": 90.0}75{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_545.wav", "doc_id": "dvGkKzmIaN.seg_545", "src_text": "Welcome to discuss with us.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir kommen, um mit Ihnen zu diskutieren.", "score": 65.0}76{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_831.wav", "doc_id": "GvEBWkLmuI.seg_831", "src_text": "However, these measures have various limitations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Maßnahmen haben jedoch verschiedene Einschränkungen:", "score": 98.0}77{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_100.wav", "doc_id": "uZBWfYjYnf.seg_100", "src_text": "So what is our solution?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Was ist die Lösung?", "score": 85.0}78{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_180.wav", "doc_id": "SLpqvupgvW.seg_180", "src_text": "Which is the alternative question.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das ist die alternative", "score": 93.0}79{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_733.wav", "doc_id": "XejEJmgUmE.seg_733", "src_text": "So the minimal pair paradigm basically evaluates language models on top of acceptability judgments.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Das Minimalpaar-Paradigma bewertet Sprachmodelle im Wesentlichen auf der Grundlage von Akzeptabilitätsurteilen,", "score": 89.0}80{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_98.wav", "doc_id": "uZBWfYjYnf.seg_98", "src_text": "And training and maintaining several models to reach different latency regimes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und das Training und Wartung mehrerer Modelle, um unterschiedliche Latenzregime zu erreichen,", "score": 89.0}81{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_858.wav", "doc_id": "GvEBWkLmuI.seg_858", "src_text": "So, really just only the positive or at least non-negative ones.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Also sind es wirklich nur die positiven oder zumindest nicht negativen, und", "score": 99.0}82{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_71.wav", "doc_id": "TVCREhgqUP.seg_71", "src_text": "That's why in the second step we use another model to predict a permutation to put them into the right order.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Deshalb verwenden wir im zweiten Schritt ein anderes Modell, um die Permutation vorherzusagen und sie in die richtige Reihenfolge zu bringen.", "score": 100.0}83{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_135.wav", "doc_id": "wLqFAuDnKa.seg_135", "src_text": "The difference observed is of more than one BLEURT points.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "– zeigt einen Unterschied von mehr als einem Blurred", "score": 62.0}84{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_557.wav", "doc_id": "rISrKoXQCx.seg_557", "src_text": "This ensures us to do automatic evaluation well grounded in political science literature.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "der politische Komplextest, um sicherzustellen, dass wir automatische Bewertungen gewährt werden, die in der politischen Wissenschaft", "score": 80.0}85{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_854.wav", "doc_id": "GvEBWkLmuI.seg_854", "src_text": "Now for some results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Nun, die ersten Stereotypen,", "score": 6.0}86{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_450.wav", "doc_id": "hgIDlKNiFM.seg_450", "src_text": "In total, we have seven models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wobei eine Zeile eine Hundertachtundzwanzig-Gigabyte-Zeile, eine Zeile", "score": 100.0}87{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_796.wav", "doc_id": "WTTtiRKFZI.seg_796", "src_text": "\"Marge read this absolutely fascinating book about bees yesterday.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Sätze gut, Marcherds Buch über die Biscuits von gestern", "score": 40.0}88{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_340.wav", "doc_id": "dJGfOSFgZO.seg_340", "src_text": "Thank you for watching.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wie sich die konversationelle KI weiterentwickelt.", "score": 0.0}89{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_608.wav", "doc_id": "oeooqChmKK.seg_608", "src_text": "And second, background knowledge such as \"Judges decide cases in law courts.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Kenntnis, dass ein Bediensteter ein Richter ist, und zweitens Hintergrundwissen, z. B. die Kenntnis, dass Richter Fälle in Gerichtshöfen entscheiden.", "score": 90.0}90{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_13.wav", "doc_id": "aQpIWggfCo.seg_13", "src_text": "We sample 100 specific goals and evaluate the scripts generated from large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir nehmen einhundert spezifische Ziele und bewerten die aus größeren Modellen generierten Skripte.", "score": 60.0}91{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_389.wav", "doc_id": "WBLMIsdIrq.seg_389", "src_text": "But if the previous sentence was \"Could it be anything serious, doctor?\", then \"mole\" refers to a birthmark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "aber wenn der vorherige Satz „Könnte es etwas Ernstes sein, Doktor?“ lautete, dann bezieht sich Mo auf ein Geburtszeichen.", "score": 60.0}92{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_765.wav", "doc_id": "XejEJmgUmE.seg_765", "src_text": "Basically, we find that the models are sensitive to the perturbed sentences in similar ways.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "stellen wir fest, dass die Modelle auf ähnliche Weise empfindlich gegenüber den periphrastischen Sätzen", "score": 60.0}93{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_544.wav", "doc_id": "dvGkKzmIaN.seg_544", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dank.", "score": 100.0}94{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_572.wav", "doc_id": "rISrKoXQCx.seg_572", "src_text": "For example, for hate speech detection, left-leaning language models are better at detecting hate speech targeting socially minority groups, however are worse at detecting hate speech targeting more powerful groups in our society.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Ergebnisse liefern. Bei der Erkennung von Hassreden, die sich auf soziale Minderheiten beziehen. Allerdings zielen wir mit unserer Hetzrede auf mehr mächtige Gruppen in unserer Gesellschaft ab.", "score": 50.0}95{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_359.wav", "doc_id": "gGbuDbHhyc.seg_359", "src_text": "First, we find that, interestingly, recent WSL methods indeed require clean validation samples to work properly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Erstens finden wir heraus, dass die neuesten WS-L-Methode tatsächlich saubere Validierungsmuster benötigt, um richtig zu funktionieren:", "score": 95.0}96{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_576.wav", "doc_id": "rISrKoXQCx.seg_576", "src_text": "There are a bunch of more examples in the appendix to further highlight that this indicates that there is a fairness issue that is very pressing regarding the political biases of language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "liefern. Es gibt noch mehr Beispiele in den Anhängen, um zu betonen, dass dies ein Fairnessproblem ist, das sehr dringend ist, was die politischen Vorurteile betrifft. Sprachmodelle, beispielsweise, wenn rechte Sprachmodelle darauf", "score": 65.0}97{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_102.wav", "doc_id": "uZBWfYjYnf.seg_102", "src_text": "Use only one model for every latency regime and handle latency through specific parameters.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "nur ein Modell für jedes Latenzregime und handhaben Sie die Latenz über spezifische Parameter.", "score": 100.0}98{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_502.wav", "doc_id": "dvGkKzmIaN.seg_502", "src_text": "Are you copying my model?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Werbevideo über Papier zu zeigen:", "score": 70.0}99{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_370.wav", "doc_id": "gGbuDbHhyc.seg_370", "src_text": "However, if we allow to continue fine-tuning on the clean samples, then FTw performs equally well as other methods.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wenn wir jedoch die feine Abstimmung der ausgewählten Proben fortsetzen dürfen, dann funktioniert FTW genauso gut wie andere Methoden.", "score": 70.0}100{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_670.wav", "doc_id": "FLkGnzVRew.seg_670", "src_text": "We also find that iterative update is useful for transfer learning from a different domain, whereas in domain active annotations benefit from cumulative update.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir stellen auch fest, dass die iterative Aktualisierung nützlich ist, um von einem anderen Domänengebiet zu lernen, wobei aktive Anmerkungen im Domänengebiet von kumulativen Aktualisierungen profitieren.", "score": 100.0}101{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_632.wav", "doc_id": "FLkGnzVRew.seg_632", "src_text": "Hello, my name is Vasudha and I'm a Computer Science PhD candidate at Stony Brook University.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hallo, mein Name ist Vasudha und ich bin Doktorand für Informatik an der Stony Brook University.", "score": 100.0}102{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_7.wav", "doc_id": "aQpIWggfCo.seg_7", "src_text": "In this paper, we define the problem of constrained language planning which imposes different constraints on the goals of planning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In diesem Papier definieren wir das Problem der eingeschränkten Sprachplanung. Dies setzt unterschiedliche Einschränkungen auf die Ziele der Planung; ein", "score": 98.0}103{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_298.wav", "doc_id": "PIZEXUFLAR.seg_298", "src_text": "We use one instruction versus 5 instruction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wobei wir eine oder fünf Anweisungen verwenden, da wir", "score": 85.0}104{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_861.wav", "doc_id": "GvEBWkLmuI.seg_861", "src_text": "In our analysis, we reveal how these seemingly positive portrayals reflect harmful patterns.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In unserer Analyse werden wir darlegen, wie diese scheinbar positiven Porträts schädliche Muster reflektieren.", "score": 90.0}105{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_94.wav", "doc_id": "uZBWfYjYnf.seg_94", "src_text": "Simultaneous speech translation, or SimulST, is the process of translating spoken language into a text in another language in real time, enabling cross-language communication.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "oder Simultandolmetschen oder Simultandolmetschen?", "score": 50.0}106{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_428.wav", "doc_id": "WBLMIsdIrq.seg_428", "src_text": "To summarize, we perform a data-driven analysis across 14 language pairs to identify when translations require context and then we use our findings to build a benchmark for document-level machine translation which can help us identify which discourse phenomena models can handle well or not, and which translation systems are good at document-level translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zusammenfassend führen wir Datengetriebene Analysen in vierzehn Sprachen durch, um eine Übersetzung zu identifizieren. Und dann können wir die Ergebnisse für die Dokumentenebene der Maschinenübersetzung verwenden, die hilfreich sein können, um zu bestimmen, ob sich die Modelle auf die Dokumentenebene der Maschinenübersetzung beziehen oder nicht, und ob sich die Übersetzungs-Systeme auf die Dokumentenebene der Maschinenübersetzung beziehen.", "score": 60.0}107{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_476.wav", "doc_id": "SUkmfOTvGi.seg_476", "src_text": "So what is needed for a good generalization?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Also, was ist für eine gute Verallgemeinerung erforderlich? ch", "score": 65.0}108{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_664.wav", "doc_id": "FLkGnzVRew.seg_664", "src_text": "Note that the performance is significantly lower for random.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "beachten Sie, dass die Leistung für Zufall deutlich niedriger ist.", "score": 95.0}109{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_27.wav", "doc_id": "aQpIWggfCo.seg_27", "src_text": "We only keep the script if the target goal scores the highest in the goal set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir behalten das Skript nur bei, wenn die Ziel-Go-Punktzahl am höchsten im Zielbereich ist.", "score": 60.0}110{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_845.wav", "doc_id": "GvEBWkLmuI.seg_845", "src_text": "And also this enables direct comparison between our generated personas and the human written responses.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "auch die direkte Vergleichsmöglichkeit zwischen unseren generierten Personen und den menschlichen Antworten. Das", "score": 60.0}111{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_688.wav", "doc_id": "oaOHnMCwad.seg_688", "src_text": "And we're not trying to say that models themselves in data sets themselves have demographic identities and life experiences, but they do aggregate judgments and opinions of real people, and can thus represent certain positionalities over others.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir werden nicht versuchen zu sagen, dass Modelle in den Köpfen und Daten in ihren Köpfen demographische Identitäten und Erfahrungen im Leben haben, aber die eigenen Modelle und Daten können ihre eigenen Urteile und Meinungen über andere Menschen aggregieren und können somit bestimmte Positionen über andere darstellen.", "score": 60.0}112{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_88.wav", "doc_id": "TVCREhgqUP.seg_88", "src_text": "Our permutation method is very flexible, but it brings the challenge that finding the highest-scoring permutation is NP-hard.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Unsere Permutationsmethode ist sehr flexibel, aber sie bringt die Herausforderung mit sich, dass die Suche nach der Permutation mit der höchsten Punktzahl sehr schwer", "score": 84.0}113{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_491.wav", "doc_id": "SUkmfOTvGi.seg_491", "src_text": "For temporal drift, we did an experiment to retrain or continue to pre-train some models with more recent data and we found that the performance degrades with larger temporal gap and this confirms our hypothesis that the main cause of the performance drop is temporal drift.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "die zeitliche Drift haben wir ein Experiment durchgeführt, bei dem wir einige Modelle mit neueren Daten neu trainiert oder weiter vortrainiert haben, und wir fanden heraus, dass die Leistung mit größerem zeitlichem Abstand abnimmt. Und dies bestätigt unsere Hypothese, dass die Hauptursache des Leistungsabfalls der zeitliche Drift ist..", "score": 90.0}114{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_172.wav", "doc_id": "SLpqvupgvW.seg_172", "src_text": "This is an important problem in conversational systems and also for benchmarking LLMs' entity understanding.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies ist ein wichtiges Problem in konventionellen Systemen und auch für die Benchmarking von LLMs. Wir sind uns", "score": 85.0}115{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_387.wav", "doc_id": "WBLMIsdIrq.seg_387", "src_text": "For example, how would we translate \"mole\" in this sentence?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zum Beispiel, wie würden wir diesen Satz übersetzen?", "score": 60.0}116{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_435.wav", "doc_id": "hgIDlKNiFM.seg_435", "src_text": "We also introduced a comparison of models with multiple pre-training settings and data sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir führen auch eine Vergleichsuntersuchung von Modellen mit mehreren tretorischen Einstellungen und Datenquellen ein,", "score": 80.0}117{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_255.wav", "doc_id": "oYCKgTzTDy.seg_255", "src_text": "For example, Encoder-Decoder outperforms previous work or achieves comparable results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Beispiel übertrifft Encoder-Decoder die Fortschritte von Progress Work oder erreicht vergleichbare Ergebnisse.", "score": 65.0}118{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_13.wav", "doc_id": "aQpIWggfCo.seg_13", "src_text": "We sample 100 specific goals and evaluate the scripts generated from large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir stellten hundert spezifische Ziele vor und bewerteten die von größeren Modellen erzeugten Skripte.", "score": 60.0}119{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_159.wav", "doc_id": "SLpqvupgvW.seg_159", "src_text": "Hi!", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hi,", "score": 100.0}120{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_40.wav", "doc_id": "aQpIWggfCo.seg_40", "src_text": "We find that T5 fine-tuned on CoScript can generate scripts of higher quality than most large language models, indicating that smaller models can surpass larger models when properly trained on suitable datasets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir stellen fest, dass T5, Fine-Tuning auf CoScript, Skripte von höherer Qualität als die meisten großen Sprachmodelle generieren kann, was darauf hindeutet, dass kleinere Modelle größere Modelle unterstützen können, wenn sie auf geeigneten Datensätzen ordnungsgemäß trainiert werden.", "score": 60.0}121{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_517.wav", "doc_id": "dvGkKzmIaN.seg_517", "src_text": "However, this method either not applicable to embedding as services or lack of transferability.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "werden. jedoch ist diese Methode entweder nicht auf die Implementierung von ADS-Diensten anwendbar oder fehlt die Übertragbarkeit: daher", "score": 60.0}122{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_33.wav", "doc_id": "aQpIWggfCo.seg_33", "src_text": "Thus, we follow the idea of symbolic knowledge distillation, to distil constrained language planning datasets from large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "So folgen wir der Idee der symbolischen Wissensdestillation, um eingeschränkte Sprachplanungsdaten aus Großsprachenmodellen zu destillieren.", "score": 75.0}123{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_618.wav", "doc_id": "oeooqChmKK.seg_618", "src_text": "In the Background-Pretrain setting, we assume that the background knowledge \"Politicians seek elected seats in government\" is contained in the pretrained parameters and in inference-time context we provide the entity-specific knowledge \"Chichester is a politician.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In der Vorbereitung nehmen wir an, dass das Hintergrundwissen der Politiker, welche Mandate in der Regierung anstreben, in den Vorbereitungsparametern enthalten ist. In der Hintergrund- und Vordergrundinformation über Politiker im Kontext des Einflusses bieten", "score": 60.0}124{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_539.wav", "doc_id": "dvGkKzmIaN.seg_539", "src_text": "The results on four data sets show that our embedding marker can have great detection performance while keep great utility for downstream tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Ergebnisse auf vier Datensätzen zeigen, dass unser eingebetteter Marker eine großartige Erkennungsleistung haben kann, während er gleichzeitig eine großartige Funktionalität für nachgelagerte Aufgaben beibehält.", "score": 96.0}125{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_154.wav", "doc_id": "wLqFAuDnKa.seg_154", "src_text": "So, it seems that PaLM chooses to produce a better-sounding translation, sometimes by dropping parts of the source sentence that are made in translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Palm sich dafür entscheidet, eine bessere Übersetzung zu produzieren, indem er manchmal Teile der Quatschen in der Übersetzung weglässt.", "score": 82.0}126{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_787.wav", "doc_id": "WTTtiRKFZI.seg_787", "src_text": "The argument is based on the principle of dependency length minimization that I will explain on the basis of these examples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "das Argument basiert auf dem Prinzip der linearen Abhängigkeitsminimierung, das anhand dieser Beispiele erklärt wird.", "score": 89.0}127{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_735.wav", "doc_id": "XejEJmgUmE.seg_735", "src_text": "And in this, minimal pair paradigm, the typical way to evaluate language models is that you show like an acceptable sentence or a grammatical sentence and then you show an acceptable sentence or an ungrammatical sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und in diesem minimalen Paradigma ist die übliche Art und Weise, Sprachmodelle zu bewerten, dass Sie eine akzeptable oder grammatikalische Sätze zeigen und dann unakzeptable oder ungrammatische Sätze. Und dann ist die", "score": 81.0}128{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_461.wav", "doc_id": "hgIDlKNiFM.seg_461", "src_text": "All the pre-trained model obtained from NACHOS are freely available on Hugging Face, and under the MIT license, and all the training scripts are on our GitHub repository.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "das vorgebildete Modell, das wir von Natcos erhalten haben, sind kostenlos verfügbar und alle Trainingsdaten sind in unserem GitHub-Repository. Frau Präsidentin, wir", "score": 53.0}129{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_330.wav", "doc_id": "dJGfOSFgZO.seg_330", "src_text": "You can see how the combination of all ABC-Eval metrics explains over 25% of conversation quality, and as you remove the metrics one at a time, most of them result in losing a decent amount of information about the quality.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "sehen, wie die Kombination aller ABC-Eval-Metriken über fünfzig Prozent der Gesprächsqualität erklärt, und wenn man die Messungen einmal entfernt, verliert man den größten Teil der Informationen über die Qualität. Auf", "score": 66.0}130{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_511.wav", "doc_id": "dvGkKzmIaN.seg_511", "src_text": "The watermark method need to meet the following properties.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die Wasserzeichenmethode muss die folgenden Eigenschaften erfüllen:", "score": 99.0}131{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_407.wav", "doc_id": "WBLMIsdIrq.seg_407", "src_text": "And this can be explained because English doesn't have dual pronouns, so you need context to determine if a pronoun is dual when translating into Arabic.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "und dies kann erklärt werden, weil Englisch keine Pronomen hat, die in Arabisch übersetzt werden", "score": 35.0}132{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_273.wav", "doc_id": "PIZEXUFLAR.seg_273", "src_text": "These tasks are derived from 21 existing open-source dataset and each task is equipped with five expert written instructions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Aufgaben werden aus einundzwanzig bestehenden Datensätzen mit offenen Quellen abgeleitet, und jede Aufgabe ist mit fünf ausgeschriebenen Anweisungen ausgestattet.", "score": 82.0}133{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_611.wav", "doc_id": "oeooqChmKK.seg_611", "src_text": "We have defined three settings of KITMUS.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir haben drei Einstellungen von Kidmos definiert.", "score": 87.0}134{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_699.wav", "doc_id": "oaOHnMCwad.seg_699", "src_text": "In Live in the Wild is an online experimentation platform where we can recruit divers volunteers.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Lab in the Wild ist eine Online-Experimentierplattform, auf der wir verschiedene Freiwillige rekrutieren können,", "score": 92.0}135{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_442.wav", "doc_id": "hgIDlKNiFM.seg_442", "src_text": "So we ask ourselves a question about what is the most appropriate data sources for a wide range of usage and those crawled data are good substitution for clinical data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "was die am besten geeigneten Datenquellen für einen breiten Anwendungsbereich sind, und diese Daten sind eine gute Alternative für klinische Daten.", "score": 94.0}136{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_369.wav", "doc_id": "gGbuDbHhyc.seg_369", "src_text": "As we can see from the figures, the vanilla model, termed FTw, initially underperforms more complicated WSL methods, like COSINE.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wie wir aus den Zahlen sehen können, unterperformt das Valina-Modell, das FTVW genannt wird, zunächst gegenüber komplexeren WSL-Methoden wie CoSine.", "score": 85.0}137{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_445.wav", "doc_id": "hgIDlKNiFM.seg_445", "src_text": "Is it 4 gigabytes, 8 gigabytes, or more?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Ist es 4 GB, 8 GB oder mehr?", "score": 100.0}138{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_538.wav", "doc_id": "dvGkKzmIaN.seg_538", "src_text": "We assume the provider apply wiki text data set to count word frequency.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir nehmen an, dass der Anbieter den Datensatz WikiText anwendet, um die Wortfrequenz zu ermitteln.", "score": 75.0}139{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_667.wav", "doc_id": "FLkGnzVRew.seg_667", "src_text": "We find that PRC has the highest percentage of dissonance and works best for rare class.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "stellen fest, dass P.R.C. der höchste Prozentsatz an Unterschieden und Arbeiten für die Klasse ist,", "score": 59.0}140{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_501.wav", "doc_id": "dvGkKzmIaN.seg_501", "src_text": "It's my pleasure to give a short advertisement video of our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Es ist mir eine Freude, ein kurzes", "score": 75.0}141{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_852.wav", "doc_id": "GvEBWkLmuI.seg_852", "src_text": "So in our method, we first designate what the unmarked and marked groups are, and then we compare the personas using the Fightin’ Words method, which is basically using weighted log-odds ratios to distinguish the top words for each marked group.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Mit unserer Methode werden wir zunächst beschreiben, was die unmarkierten und markierten Gruppen sind. Und wir vergleichen die Personen, die die Kampfwörtermethode verwenden, die im Grunde Gewichtsverhältnisse verwendet, um die Top-Wörter für jede Markgruppe zu unterscheiden.", "score": 60.0}142{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_758.wav", "doc_id": "XejEJmgUmE.seg_758", "src_text": "So here we are choosing or creating sentences from acceptable and unacceptable domains from the same BLiMP or SyntaxGym dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hier wählen oder erstellen wir Sätze aus akzeptablen und inakzeptablen Domänen aus dem gleichen Blip- oder Syntax-Datensatz.", "score": 60.0}143{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_702.wav", "doc_id": "oaOHnMCwad.seg_702", "src_text": "Afterwards to stay engaged in the study, they can compare their responses to an AI and others.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Danach, um sich in der Studie zu engagieren, können sie ihre Antworten mit einer AI und anderen vergleichen.", "score": 80.0}144{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_197.wav", "doc_id": "SLpqvupgvW.seg_197", "src_text": "For songs, we simply show a Google search link to each song and then ask the annotators to listen to at least some of each song, and read about each song.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "den Zwanzigern zeigen, für Songs zeigen wir einfach einen Google-Suchlink. Und dann bitten Sie die Annotatoren, zumindest einen Teil des Liedes zu hören und über das Lied zu lesen.", "score": 90.0}145{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_705.wav", "doc_id": "oaOHnMCwad.seg_705", "src_text": "We then compared these annotations with Dynahate, Perspective API, Rewire API, Hate Roberta and GPT 4.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "dann diese Anmerkungen mit DynaHate, Perspective API, Rewire API, Hate Roberta und GPT-4. Art und", "score": 96.0}146{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_482.wav", "doc_id": "SUkmfOTvGi.seg_482", "src_text": "And last but not least, we all know that the number of fine tuning examples directly affects the performance of a downstream task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und nicht zuletzt wissen wir alle, dass die Anzahl der Feintuning-Beispiele die Leistung einer Downstream-Aufgabe direkt beeinflusst.", "score": 96.0}147{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_576.wav", "doc_id": "rISrKoXQCx.seg_576", "src_text": "There are a bunch of more examples in the appendix to further highlight that this indicates that there is a fairness issue that is very pressing regarding the political biases of language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "anzuzeigen. Es gibt noch viele weitere Beispiele in den Anhängen. Dies deutet darauf hin, dass es sich um ein Fairnessproblem handelt, das in Bezug auf die politischen Sprachmodelle sehr dringend ist.", "score": 65.0}148{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_590.wav", "doc_id": "oeooqChmKK.seg_590", "src_text": "This work is a collaboration between McGill University, Mila, and Microsoft Research.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese Arbeit ist eine Zusammenarbeit zwischen der Universität Mila und Microsoft Research.", "score": 90.0}149{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_230.wav", "doc_id": "oYCKgTzTDy.seg_230", "src_text": "We use Google Translate API to translate source to the target language, then use monolingual model to train and evaluation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir verwenden Google Translate API, um die Quelle in die Ziel-Sprache zu übersetzen, und dann verwenden wir ein monolinguales Modell, um zu trainieren und zu evaluieren.", "score": 100.0}150{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_253.wav", "doc_id": "oYCKgTzTDy.seg_253", "src_text": "We found that, by comparing the green and orange line, we found the Zero-shot setting, the Cross-lingual transfer performance gap is significant, and then comparing the blue and orange lines, we found that with the Few-shot setting the transfer gap is shortened rapidly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Einzelsprachübertragung. Wir stellten fest, dass sich die Übertragungsleistung für die Einstellung „grün-orange“ erheblich verringert, während sich die Übertragungsleistung für die Einstellung „blau-orange“ mit wenigen Schüssen erheblich verringert.", "score": 88.0}151{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_864.wav", "doc_id": "GvEBWkLmuI.seg_864", "src_text": "This contributes to a long legacy of discrimination and othering for these groups.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies trägt zu einer langen Geschichte von Diskriminierung und anderen Dingen für diese Gruppen bei.", "score": 60.0}152{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_695.wav", "doc_id": "oaOHnMCwad.seg_695", "src_text": "And we ought to do this over looking at the demographics of original data sets annotators, because, usually only a few annotators annotate each instance and because demographics are rarely collected and shared.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wiederauswerten von Daten mit diversen Annotatoren, und wir möchten dies über die Demografien der ursprünglichen Datenmengen Annotatoren tun, weil normalerweise nur wenige Annotatoren jede Instanz annotieren und weil Demografien seltener gesammelt und geteilt", "score": 60.0}153{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_313.wav", "doc_id": "dJGfOSFgZO.seg_313", "src_text": "The common practice is to use human evaluation, such as by asking human judges to select which of two conversations is better or to rate conversations given a Likert scale.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die allgemeine Praxis ist, menschliche Bewertungen zu verwenden, beispielsweise, indem man menschliche Richter bittet, zu entscheiden, welche zwei Gespräche besser sind, oder indem man Gespräche, die eine niedrige Bewertung erhalten, mit einer niedrigeren Bewertung bewertet.", "score": 60.0}154{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_86.wav", "doc_id": "TVCREhgqUP.seg_86", "src_text": "In addition, sometimes there are multiple permutations that are consistent with the data, but the linguistically correct one is latent.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Darüber hinaus gibt es manchmal mehrere Permutationen, die mit den Daten übereinstimmen, aber die sprachlich korrekte ist latent.", "score": 100.0}155{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_493.wav", "doc_id": "SUkmfOTvGi.seg_493", "src_text": "And these goes hand in hand, we can't just have one ingredient but throw out the others.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "diese Ziele gehen Hand in Hand, wir können nicht nur ein Zutaten haben, sondern alle anderen durchgehen.", "score": 60.0}156{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_164.wav", "doc_id": "SLpqvupgvW.seg_164", "src_text": "\"Did you mean 'Easy on Me' or 'I Gotta Feeling'?\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Frage: Haben Sie „einfach“ oder „Gefühl“ gemeint? Der", "score": 50.0}157{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_741.wav", "doc_id": "XejEJmgUmE.seg_741", "src_text": "So that is the approach.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das ist der Ansatz,", "score": 95.0}158{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_150.wav", "doc_id": "wLqFAuDnKa.seg_150", "src_text": "But, PaLM comes pretty close to a commercial system.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Pan-Übersetzungen. Dann? 2 kommt unserem kommerziellen System ziemlich nahe:", "score": 50.0}159{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_597.wav", "doc_id": "oeooqChmKK.seg_597", "src_text": "In this work, we propose a diagnostic test suite for knowledge integration.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In dieser Arbeit schlagen wir ein diagnostisches Testsystem für die Wissensintegration vor.", "score": 100.0}160{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_690.wav", "doc_id": "oaOHnMCwad.seg_690", "src_text": "However these works really don't look at comparing end users with the datasets and models themselves, and studying model and data set positionality is increasingly important as NLP tasks become more subjective and socially oriented, and it's challenging to characterise how these positionalities are skewed because not all decisions are documented and many models are hidden behind APIs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Arbeiten vergleichen jedoch nicht wirklich Endnutzer mit den Datensätzen und Modellen selbst. Das Studium der Modell- und Datenpositionalität wird immer wichtiger, da die NP-Tests subjektiver und sozial orientierter werden. Es ist schwierig, zu charakterisieren, wie diese Positionalitäten verzerrt sind, weil nicht alle Entscheidungen dokumentiert sind und viele Modelle hinter APIs versteckt sind.", "score": 60.0}161{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_177.wav", "doc_id": "SLpqvupgvW.seg_177", "src_text": "In the first bubble, Bob says, \"Remember that song we were listening to yesterday?\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In dem ersten Bläser sagt Bob: „Erinnere dich an das Lied, das wir gestern gehört haben“,", "score": 60.0}162{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_366.wav", "doc_id": "gGbuDbHhyc.seg_366", "src_text": "The right figure shows the performance difference between fine-tuning approaches, which are directly applied on the clean data, and WSL approaches, which use the clean data for validation only.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die rote Abbildung zeigt den Leistungsunterschied zwischen Fine-Tuning-Ansätzen, die direkt auf sauberen Daten angewendet werden, und WSL-Ansätzen, die nur für die Validierung von sauberen Daten verwendet werden.", "score": 60.0}163{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_769.wav", "doc_id": "XejEJmgUmE.seg_769", "src_text": "Please read our paper for more details of our experiments.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Bitte lesen Sie unser Papier für weitere Details zu unseren", "score": 95.0}164{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_149.wav", "doc_id": "wLqFAuDnKa.seg_149", "src_text": "Nevertheless, specialized state-of-the-art systems have a substantial advantage over the PaLM translations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dennoch haben spezialisierte Systeme einen erheblichen Vorteil gegenüber den", "score": 85.0}165{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_680.wav", "doc_id": "oaOHnMCwad.seg_680", "src_text": "But that's not really the case for Aditya Sharma.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "aber das ist nicht wirklich der Fall für Aditya Sharma, wo", "score": 87.0}166{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_389.wav", "doc_id": "WBLMIsdIrq.seg_389", "src_text": "But if the previous sentence was \"Could it be anything serious, doctor?\", then \"mole\" refers to a birthmark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Aber wenn die vorherige Aussage lautete, „Könnte es irgendetwas Ernstes, Doktor?“, bezieht sich „Moe“ auf einen Geburtsurkunden.", "score": 78.0}167{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_865.wav", "doc_id": "GvEBWkLmuI.seg_865", "src_text": "Furthermore, there's a lot of common tropes that are reflected in these words, especially for women of color.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Darüber hinaus sind in diesen Wörtern viele Komposita enthalten, insbesondere für Frauen", "score": 65.0}168{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_310.wav", "doc_id": "dJGfOSFgZO.seg_310", "src_text": "And today we'll tell you all about ABC-Eval, a new dimensional approach to evaluating conversational AI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und heute erzählen wir Ihnen alles über Abcevel, einen neuen dimensionalen Ansatz zur Bewertung von Konversations-AI.", "score": 85.0}169{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_361.wav", "doc_id": "gGbuDbHhyc.seg_361", "src_text": "As shown in this figure, if there are no clean validation samples, then the trained models cannot generalize beyond the original weak labels, meaning that the training is pointless.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wie in dieser Abbildung zu sehen ist, wenn es keine sauberen Validierungsmuster gibt, dann können die Trendmodelle nicht über die ursprünglichen Bit-Labels generalisiert werden. Das bedeutet, dass die Doktrin sinnlos ist.", "score": 53.0}170{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_16.wav", "doc_id": "aQpIWggfCo.seg_16", "src_text": "Then we conduct detailed analysis to investigate why learning models fail.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dann führen wir detaillierte Analysen durch, um zu untersuchen, wofür Landnutzungsmodelle geeignet sind. Die Ergebnisse", "score": 59.0}171{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_435.wav", "doc_id": "hgIDlKNiFM.seg_435", "src_text": "We also introduced a comparison of models with multiple pre-training settings and data sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir stellen auch einen Vergleich von Modellen mit multiplen prädiktiven Einstellungen und Datenquellen an,", "score": 90.0}172{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_597.wav", "doc_id": "oeooqChmKK.seg_597", "src_text": "In this work, we propose a diagnostic test suite for knowledge integration.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In diesem Projekt schlagen wir einen diagnostischen Test vor, um die Wissensintegration zu ermöglichen.", "score": 90.0}173{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_243.wav", "doc_id": "oYCKgTzTDy.seg_243", "src_text": "And, we also evaluate Encoder-Decoder models, which is Multilingual Pretrained Encoder-Decoder Models, such as mBART and mT5.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "steht. Und wir bewerten auch Encoder-Decoder-Modelle, die multilinguale trainierte Encoder-Decoder-Modelle sind, wie z. B. Anbert und M.T.5.", "score": 53.0}174{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_164.wav", "doc_id": "SLpqvupgvW.seg_164", "src_text": "\"Did you mean 'Easy on Me' or 'I Gotta Feeling'?\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Frage: Wollen Sie es mir leicht machen oder haben Sie ein Gefühl dafür?", "score": 0.0}175{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_210.wav", "doc_id": "SLpqvupgvW.seg_210", "src_text": "For example, when the language model retrieves the background knowledge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wenn das Sprachmodell das Hintergrundwissen zurückgibt.", "score": 90.0}176{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_239.wav", "doc_id": "oYCKgTzTDy.seg_239", "src_text": "We train on one source language and transfer to another language.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "einer Quellsprache und einer Ziel-Sprache.", "score": 40.0}177{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_374.wav", "doc_id": "gGbuDbHhyc.seg_374", "src_text": "Our concrete recommendations for future work are as follows.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Unsere konkreten Empfehlungen für zukünftige Arbeiten sind wie folgt.", "score": 100.0}178{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_8.wav", "doc_id": "aQpIWggfCo.seg_8", "src_text": "An abstract goal can be inherited by different real-life specific goals with multi-faceted constraints.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Ein abstraktes Ziel kann durch unterschiedliche realleben-spezifische Ziele mit multifaktoriellen Einschränkungen vererbt werden.", "score": 94.0}179{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_862.wav", "doc_id": "GvEBWkLmuI.seg_862", "src_text": "First, from our groups, the top words include things like \"culture\", \"tradition\", \"proud\", and \"exotic\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zunächst für Markgruppen: Die oberen Wörter beinhalten Dinge wie Kultur, Tradition, stolz und exotisch.", "score": 86.0}180{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_198.wav", "doc_id": "SLpqvupgvW.seg_198", "src_text": "Here's for example, the Google search result for the song \"Easy on Me.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hier ist beispielsweise das Google-Suchergebnis für das Lied „Easy“. Für die", "score": 80.0}181{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_103.wav", "doc_id": "uZBWfYjYnf.seg_103", "src_text": "And leverage the knowledge already acquired by the model through the attention mechanism between audio input and textual output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und nutzen Sie die bereits durch das Modell erworbenen Kenntnisse durch die Aufmerksamkeit zwischen Audio-Eingabe und Text. Output, also", "score": 80.0}182{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_168.wav", "doc_id": "SLpqvupgvW.seg_168", "src_text": "This could happen when the user cannot remember the name of the song.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "dies könnte passieren, wenn der Benutzer sich den Namen des Geräts nicht erinnern kann.", "score": 82.0}183{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_761.wav", "doc_id": "XejEJmgUmE.seg_761", "src_text": "Now this and this is very large like this effect, increases throughout the context length and this would probably affect like newer language models which has large context window.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Jetzt ist dies sehr groß – wie dieser Effekt sich über die Kontextlänge erstreckt und dies wahrscheinlich neue Sprachmodelle mit großen Kontextfenstern beeinflusst.", "score": 77.0}184{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_534.wav", "doc_id": "dvGkKzmIaN.seg_534", "src_text": "The cosine and L2 similarity between the requested embedding and the target embedding are computed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die Koinosität und L2-Similität zwischen dem angeforderten Einbetten und dem Ziel-Einbetten werden berechnet;", "score": 80.0}185{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_516.wav", "doc_id": "dvGkKzmIaN.seg_516", "src_text": "Existing works can be broadly classified into four categories.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Bestehende Werke können im Großen und Ganzen in vier Kategorien eingeteilt werden.", "score": 98.0}186{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_248.wav", "doc_id": "oYCKgTzTDy.seg_248", "src_text": "I think this is known as the \"Curse of Multilinguality\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "denke, dies wird als Fluch der Mehrsprachigkeit bezeichnet.", "score": 96.0}187{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_96.wav", "doc_id": "uZBWfYjYnf.seg_96", "src_text": "Specific architectures are usually trained, introducing additional modules to be optimized.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Spezifische Architekturen werden üblicherweise trainiert, um zusätzliche Module einzuführen, die optimiert werden können. Langfristige,", "score": 94.0}188{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_541.wav", "doc_id": "dvGkKzmIaN.seg_541", "src_text": "The legend of the figures means the number of triggers in each sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wobei die Legende der Figuren die Anzahl der Auslöser in jedem Satz bedeutet.", "score": 80.0}189{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_596.wav", "doc_id": "oeooqChmKK.seg_596", "src_text": "Therefore, successful models for knowledge-intensive NLU tasks require the ability to integrate and use both pretrain-time and inference-time knowledge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Deshalb erfordern erfolgreiche Modelle für N-LU-Aufgaben die Fähigkeit, Vor-Trainingszeit und Inferenzzeit-Wissen zu integrieren und zu verwenden.", "score": 95.0}190{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_526.wav", "doc_id": "dvGkKzmIaN.seg_526", "src_text": "When a user send a sentence to the provider service the provider counts the trigger number in the sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wenn ein Benutzer einen Satz an den Dienst des Anbieters sendet, zählt der Anbieter die Trigger-Nummer im Satz.", "score": 100.0}191{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_139.wav", "doc_id": "wLqFAuDnKa.seg_139", "src_text": "So in this example here, where we perform translation from German into English, the German sentences, the source sentences, are marked with German colon and the English translations with English colon.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In diesem Beispiel hier, wo wir Übersetzungen von Deutsch ins Englische durchführen, markieren wir die deutschen Sätze. Diese Sätze sind mit einem deutschen Kolon markiert und die englischen Übersetzungen mit englischen Spalten.", "score": 95.0}192{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_278.wav", "doc_id": "PIZEXUFLAR.seg_278", "src_text": "In which the input text, images, instructions and bounding boxes are represented in the same token space.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "in der der Eingabetext, Bilder, Anweisungen und Bindungskästchen in derselben Token-Raum repräsentiert werden.", "score": 97.0}193{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_545.wav", "doc_id": "dvGkKzmIaN.seg_545", "src_text": "Welcome to discuss with us.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir werden uns mit Ihnen unterhalten.", "score": 0.0}194{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_445.wav", "doc_id": "hgIDlKNiFM.seg_445", "src_text": "Is it 4 gigabytes, 8 gigabytes, or more?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Es ist vier Gigabyte, acht Gigabyte oder mehr.", "score": 86.0}195{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_713.wav", "doc_id": "oaOHnMCwad.seg_713", "src_text": "So for GPT 4, in the social acceptability task, we find that it's most aligned to people with a college education or Graduate School education and we find the same for Dynahate where it's most aligned to people with a college education.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wir die meisten Zuordnungen zu Personen mit Hochschulbildung oder Hochschulabschluss. Und wir finden das Gleiche für Donieght, wo es sich hauptsächlich um Menschen mit einer Hochschulbildung handelt.", "score": 0.0}196{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_191.wav", "doc_id": "SLpqvupgvW.seg_191", "src_text": "The second one is when the entities have similar titles, for example, two books with the name \"The Return\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zufallsauswahl, die zweite Methode ist die gleichnamige Zufallsauswahl, z.B. zwei Bücher mit dem Namen 'The Return',", "score": 60.0}197{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_122.wav", "doc_id": "wLqFAuDnKa.seg_122", "src_text": "Hello everyone, my name is David Vilar, and I will be giving a short review of the paper \"Prompting PaLM for Translation: Assessing Strategies and Performance.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "1 Hallo, ich bin Irwan, mein Name ist Ayesha Vilar und ich werde eine kurze Zusammenfassung des Papiers 'Promoting Powerful Translation: Assessing Strategies and Performance' geben.", "score": 63.0}198{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_532.wav", "doc_id": "dvGkKzmIaN.seg_532", "src_text": "Back door data set contains sentences of which all words belong to the trigger set while all words in the sentences of benign data set do not belong to the trigger sets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Der Backdoor-Datensatz enthält Sätze, in denen alle Wörter zum Trigger-Set gehören, während alle Wörter im günstigen Datensatz nicht zum Trigger-Set gehören.", "score": 97.0}199{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_818.wav", "doc_id": "WTTtiRKFZI.seg_818", "src_text": "In such cases, the left conjunct prefers to be shorter; the most of the biggest difference between the two conjuncts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In solchen Fällen ist die linke Konjunktion bevorzugt, die größere Differenz zwischen den beiden Wörtern. Allerdings", "score": 60.0}200{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_871.wav", "doc_id": "GvEBWkLmuI.seg_871", "src_text": "So rather than actually working towards changing those obstacles, it puts pressure on those people to overcome them, which leads to a very negative health outcomes for these people, among other harms.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "anstatt tatsächlich daran zu arbeiten, diese Hindernisse zu ändern, und Druck auf diese Menschen auszuüben. - die zu sehr negativen Gesundheitsauswirkungen für diese Menschen und andere Schäden führt. Im Allgemeinen finden wir,", "score": 60.0}201{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_658.wav", "doc_id": "FLkGnzVRew.seg_658", "src_text": "Next, we determine the best method to update a model with new data from each round of active learning and annotations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Als nächstes werden wir die beste Methode ermitteln, um ein Modell mit neuen Daten aus jeder Runde des aktiven Lernens und der Anmerkungen zu aktualisieren.", "score": 100.0}202{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_408.wav", "doc_id": "WBLMIsdIrq.seg_408", "src_text": "And similarly, we find that certain languages also require context when we want to choose the appropriate verb form.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir eine geeignete Verbform wählen möchten.", "score": 100.0}203{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_506.wav", "doc_id": "dvGkKzmIaN.seg_506", "src_text": "Embedding as services is one of the services built upon large language models to assist various, NLP tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Embedding ADS ist eine der Dienste, die auf großen Sprachmodellen gebaut wurden, um verschiedene NLP-Aufgaben zu unterstützen.", "score": 60.0}204{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_682.wav", "doc_id": "oaOHnMCwad.seg_682", "src_text": "This is an example of a design bias where we see systematic performance differences of technology between populations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Dies ist ein Beispiel für ein Designfehler, bei dem wir systematische Leistungsunterschiede zwischen Technologien zwischen Bevölkerungsgruppen sehen.", "score": 65.0}205{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_514.wav", "doc_id": "dvGkKzmIaN.seg_514", "src_text": "Third, the watermark should be covert enough to the attacker or the attacker can remove the watermark easily.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Drittens sollte die Wassermarke ausreichend für den Angreifer abgedeckt sein, oder der Angreifer kann die Wassermarke leicht entfernen.", "score": 100.0}206{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_39.wav", "doc_id": "aQpIWggfCo.seg_39", "src_text": "With CoScript we can try smaller but specialized models for constrained language planning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "für eingeschränkte Sprachplanung. Auf", "score": 50.0}207{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_375.wav", "doc_id": "gGbuDbHhyc.seg_375", "src_text": "First, report the model selection criteria.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zunächst müssen die Kriterien für die Modellauswahl angegeben werden;", "score": 100.0}208{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_771.wav", "doc_id": "WTTtiRKFZI.seg_771", "src_text": "Hi, my name is Adam Przepiórkowski and this talk is about the Dependency Structure of Coordination.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo, mein Name ist Adam Schyrkovski, und dieses Gespräch dreht sich um die Abhängigkeitsstruktur der Koordination.", "score": 71.0}209{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_723.wav", "doc_id": "oaOHnMCwad.seg_723", "src_text": "I mean, we want to emphasise that inclusive NLP isn't just making.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Ich meine, wir möchten betonen, dass eine inklusive NLP nicht nur bedeutet, dass alle", "score": 100.0}210{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_749.wav", "doc_id": "XejEJmgUmE.seg_749", "src_text": "So here the sentences are still coming from a, relevant data sets but it's not from the same data set that you are evaluating with.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hier werden also die Sätze immer noch aus relevanten Datensätzen, aber nicht aus dem gleichen Datensatz, mit dem Sie die Bewertung durchführen, und", "score": 99.0}211{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_210.wav", "doc_id": "SLpqvupgvW.seg_210", "src_text": "For example, when the language model retrieves the background knowledge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wenn das Sprachmodell das Hintergrundwissen wiedererlangt.", "score": 75.0}212{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_773.wav", "doc_id": "WTTtiRKFZI.seg_773", "src_text": "So for example, in the universal dependencies, the structure of the coordination, Lisa, Bart, and Maggie, such that the first conjunct is the head of the whole coordinate structure.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "so zum Beispiel bei universellen Abhängigkeiten, die Struktur der koordinierten Koordination Lisa A. B. und Maggie. Es ist so, dass der erste Konjunkt ist der Kopf der gesamten Kordstruktur, also", "score": 78.0}213{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_418.wav", "doc_id": "WBLMIsdIrq.seg_418", "src_text": "We then use the MuDA tagger, by applying the tagger on a parallel corpus that we want to use for evaluation and we apply our translation metrics of choice on the context-dependent examples that the MuDA tagger has identified.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dann verwenden wir den Mudatagger, indem wir den Tag auf einen parallelen Korpus anwenden, den wir für die Bewertung verwenden möchten, und unsere Übersetzungsmetriken der Wahl auf die kontextabhängigen Beispiele, die der Mudatagger identifiziert hat.", "score": 76.0}214{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_663.wav", "doc_id": "FLkGnzVRew.seg_663", "src_text": "We find that the proposed PRC strategy works better than other state-of-the-art strategies, although the difference is small.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir stellen fest, dass die vorgeschlagene PR-Strategie besser funktioniert als andere Strategien, auch wenn der Unterschied gering ist,", "score": 91.0}215{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_768.wav", "doc_id": "XejEJmgUmE.seg_768", "src_text": "And the MPP evaluation the way that we do it currently with short and single sentence input, may not fully capture the language models abstract knowledge throughout the context window.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und die MPV-Bewertung, die Art und Weise, wie wir es derzeit korrekt mit kurzer und einziger Satz-Eingabe tun, mag nicht die abstrakte Wissen der Sprachmodelle durch den Kontextfenster", "score": 87.0}216{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_166.wav", "doc_id": "SLpqvupgvW.seg_166", "src_text": "The most obvious thing is to use a direct reference, for example by saying the name of the song \"Easy on Me\" or its position, \"the first one\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die offensichtlichste Sache ist, eine direkte Referenz zu verwenden, z.B. indem man den Namen des Liedes 'Easy on me' oder seine Position, 'Erster', sagt.", "score": 96.0}217{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_112.wav", "doc_id": "uZBWfYjYnf.seg_112", "src_text": "So we want our curves to be as high as possible on this plot.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir wollen also, dass unsere Warteschlangen auf diesem Plot so hoch wie möglich sind.", "score": 77.0}218{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_97.wav", "doc_id": "uZBWfYjYnf.seg_97", "src_text": "Long and complicated training procedures, for example, training involving different optimization objectives.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "komplizierte Trainingsverfahren, zum Beispiel das Training, das", "score": 100.0}219{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_452.wav", "doc_id": "hgIDlKNiFM.seg_452", "src_text": "These models are compared to six baseline models which are CamemBERT OSCAR 138 GB, CamemBERT OSCAR 4 GB, CamemBERT CCNET 4 GB, PubMedBERT, BioBERT, and ClinicalBERT.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "verglichen, die sich auf die folgenden Modelle beziehen: Kamnaber-Osaka 1,38 GB, Kamnaber-Osaka 4 GB, Kamnaber-Cisnet 4 GB, Plumbert-BioBERT und ClinicalBERT.", "score": 35.0}220{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_511.wav", "doc_id": "dvGkKzmIaN.seg_511", "src_text": "The watermark method need to meet the following properties.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die Wasserzeichenmethode muss die folgenden Eigenschaften erfüllen:", "score": 100.0}221{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_535.wav", "doc_id": "dvGkKzmIaN.seg_535", "src_text": "We compute the similarity difference between benign and backdoor data set which is defined as delta cosine and delta L2.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir berechnen den Unterschied der Ähnlichkeit zwischen dem Benignen und dem Hintergrund-Datensatz, der als Delta-Koinosität und Delta-L2 definiert wird.", "score": 90.0}222{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_735.wav", "doc_id": "XejEJmgUmE.seg_735", "src_text": "And in this, minimal pair paradigm, the typical way to evaluate language models is that you show like an acceptable sentence or a grammatical sentence and then you show an acceptable sentence or an ungrammatical sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und in diesem Minimalparadigma wird die typische Art, Sprachmodelle zu bewerten, so, dass man eine akzeptable oder grammatikalische Sätze zeigt und dann einen unakzeptablen oder ungrammatischen Satz zeigt. Und dann ist", "score": 82.0}223{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_325.wav", "doc_id": "dJGfOSFgZO.seg_325", "src_text": "For each of the existing methods, we collected evaluations on eight of the most commonly measured aspects of dialogue, since this is the standard practice for evaluating chat models along multiple dimensions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Für jede der vorhandenen Methoden haben wir Bewertungen zu acht der am häufigsten gemessenen Aspekte des Dialogs gesammelt, da dies die Standardpraxis für die Bewertung von Chat-Modellen in mehreren Dimensionen ist.", "score": 100.0}224{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_461.wav", "doc_id": "hgIDlKNiFM.seg_461", "src_text": "All the pre-trained model obtained from NACHOS are freely available on Hugging Face, and under the MIT license, and all the training scripts are on our GitHub repository.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "besser, aber es passt nicht gut. Das Pre-Training-Modell, das von Natasha stammt, ist frei verfügbar und auf Hugging Face sowie alle Trainingsskripte auf unserem GitHub-Repository. Also vielen", "score": 60.0}225{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_226.wav", "doc_id": "oYCKgTzTDy.seg_226", "src_text": "We provide a uniform data set XSemPLR for cross-lingual semantic parsing in multiple natural languages and meaning representations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir bieten ein uniformes Datensatz-Beispiel für die semantische Analyse von Mehrfachverbindungen in mehreren natürlichen Sprachen und Meningsrepräsentationen. Es enthält", "score": 56.0}226{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_367.wav", "doc_id": "gGbuDbHhyc.seg_367", "src_text": "As we can see, if we have 10 samples per class, direct fine-tuning starts to beat WSL approaches.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wie wir sehen können, beginnt die direkte Feinabstimmung, wenn wir zehn Proben pro Klasse haben.", "score": 64.0}227{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_388.wav", "doc_id": "WBLMIsdIrq.seg_388", "src_text": "Well, if the previous sentence was \"Things could start to get dangerous if the ministers find out\", then \"mole\" refers to a spy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Nun, wenn der vorherige Satz „Dinge könnten gefährlich werden, wenn die Minister das herausfinden“ lautet, dann bezieht sich Mo auf einen Spion.", "score": 86.0}228{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_419.wav", "doc_id": "WBLMIsdIrq.seg_419", "src_text": "And finally, we use our benchmark as well as other metrics to evaluate different models on the document-level machine translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und schließlich verwenden wir unsere Benchmarks sowie andere Metriken, um unterschiedliche Modelle zu bewerten, auf Dokumentenebene. Zunächst,", "score": 82.0}229{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_316.wav", "doc_id": "dJGfOSFgZO.seg_316", "src_text": "One approach is to simply ask human judges to evaluate several dimensions of dialogue quality, such as the relevance of model responses using existing comparative or Likert scale methods.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ein Ansatz besteht darin, einfach menschliche Urteilsfinder zu bitten, mehrere Dimensionen der Dialogqualität zu bewerten, wie z. B. die Relevanz von Modellantworten unter Verwendung von vergleichenden oder Likert-Skalenmethoden.", "score": 98.0}230{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_635.wav", "doc_id": "FLkGnzVRew.seg_635", "src_text": "Simply put, cognitive dissonance is two beliefs or actions that are inconsistent, such as this example where a person states, \"I know that cigarettes could kill me\", and then goes on to say \"I grabbed a couple of smokes after the meeting\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Einfach ausgedrückt ist kognitive Dissonanz zwei Glaubenssätze oder Handlungen, die inkonsistent sind. In diesem Beispiel, wenn eine Person sagt, ich weiß, dass die Zigaretten mich umbringen würden, und dann sage ich, ich schnorchele ein paar Zigaretten nach", "score": 62.0}231{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_488.wav", "doc_id": "SUkmfOTvGi.seg_488", "src_text": "This means that every unit of improvement that we made, on CoNLL-2003 translates to more than one unit improvement on CoNLL++ which means that there is no diminishing returns.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "hat. Dies bedeutet, dass jede Einheit der Verbesserung, die wir auf Carolo 2003 vorgenommen haben, sich zu mehr als einer Einheit Verbesserung auf Carolo + Plus übersetzt, was bedeutet, dass es keine Abnahme der Rendite gibt.", "score": 62.0}232{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_7.wav", "doc_id": "aQpIWggfCo.seg_7", "src_text": "In this paper, we define the problem of constrained language planning which imposes different constraints on the goals of planning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In diesem Papier definieren wir das Problem der eingeschränkten Sprachplanung. Dies setzt unterschiedliche Einschränkungen für die Planungsziele voraus.", "score": 80.0}233{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_153.wav", "doc_id": "wLqFAuDnKa.seg_153", "src_text": "So, in particular, the most common errors are omission errors.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "die häufigsten Fehler sind Fehler der Nichtbeachtung.", "score": 93.0}234{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_379.wav", "doc_id": "gGbuDbHhyc.seg_379", "src_text": "Finally, we have open-sourced our code.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Endlich haben wir unseren Code in der Öffentlichkeit zugänglich gemacht.", "score": 88.0}235{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_586.wav", "doc_id": "rISrKoXQCx.seg_586", "src_text": "Ok, great.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Okay, großartig,", "score": 100.0}236{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_167.wav", "doc_id": "SLpqvupgvW.seg_167", "src_text": "But sometimes an indirect reference is more appropriate to have a more natural conversation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Aber manchmal ist eine indirekte Anspielung angemessener, um eine natürlichere Konversation:", "score": 91.0}237{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_161.wav", "doc_id": "SLpqvupgvW.seg_161", "src_text": "My name is Javad Hosseini and this is a joint work with Filip Radlinski, Silvia Pareti, and Annie Louis.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Mein Name ist Jawad Hussaini, und das ist eine gemeinsame Arbeit mit Philip Radlinski, Sylvia Patry und Ani Tuis.", "score": 88.0}238{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_580.wav", "doc_id": "rISrKoXQCx.seg_580", "src_text": "We would also like to highlight that we expose the unique dilemma regarding language model political biases.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "der weiteren Diskussion möchten wir auch darauf hinweisen, dass wir die einzigartige Dilemma in Bezug auf die Sprachmodalitäten erläutern", "score": 30.0}239{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_716.wav", "doc_id": "oaOHnMCwad.seg_716", "src_text": "We find this in the GPT 4 social acceptability task as well as the Dynahate task analysis as well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir finden dies in der Gleichstellungsaufgabe. Was können wir unter Berücksichtigung der Tatsache", "score": 69.0}240{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_146.wav", "doc_id": "wLqFAuDnKa.seg_146", "src_text": "In particular, we compare the selecting prompts from the training data for the WMT evaluations on the dev data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "insbesondere vergleichen wir die. 1. Die Auswahl von Anregungen aus den Trainingsdaten der WMT-Evaluierungen oder 2. Die dev-Daten.", "score": 76.0}241{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_455.wav", "doc_id": "hgIDlKNiFM.seg_455", "src_text": "We also observe that using more data translated to better performance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und wir stellen auch fest, dass die Verwendung von mehr Daten zu besseren Leistungen führt.", "score": 99.0}242{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_704.wav", "doc_id": "oaOHnMCwad.seg_704", "src_text": "We then replicate a very similar setup for the toxicity and hate speech detection task, where they'll read an instance from Dynahate and write whether they think it's instance of hate speech.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wiederholten wir sehr ähnliche Sätze für die Toxizitäts- und Sprachdetektionsaufgabe, wobei die Fälle von „dienen“ und „rechtens“ als Beispiele für die Sprachdetektionsaufgabe dienten. Dann vergleichen wir diese Anmerkungen", "score": 19.0}243{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_56.wav", "doc_id": "TVCREhgqUP.seg_56", "src_text": "In contrast to standard machine learning evaluation, the test set does not come from the same distribution but contains structurally unseen logical forms.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Im Gegensatz zu standardmäßiger maschinellem Lernalgebra wird die Testmenge nicht aus der gleichen Verteilung stammen, sondern enthält strukturell unerkannte logische Formen.", "score": 94.0}244{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_451.wav", "doc_id": "hgIDlKNiFM.seg_451", "src_text": "To evaluate our seven models, we gather data for public and private downstream tasks such as named entity recognition, classification, part-of-speech tagging, and question answering.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Um unsere sieben Modelle zu bewerten, haben wir mehrere öffentliche und private Downstream-Tasks wie NER, Klassifikation, Part-of-Speech-Tagging, Das Modell wird mit sechs Basismodellen", "score": 69.0}245{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_822.wav", "doc_id": "WTTtiRKFZI.seg_822", "src_text": "What we see here is that when the governor is on the left, the tendency for the left conjunct to be shorter grows steadily, with the absolute difference in words, and the same is observed when there is no governor as in coordination of sentences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "dass das so ist, wenn der Gouverneur auf der Rückseite ist. Die Tendenz des linken Konjunktivs, kürzer zu werden, wächst stetig, mit dem absoluten Unterschied in den Wörtern, und dasselbe wird beobachtet,", "score": 72.0}246{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_833.wav", "doc_id": "GvEBWkLmuI.seg_833", "src_text": "Furthermore, most work in this space doesn't account for intersectionality, which is the notion that multi-faceted social identities can compound biases and be unique loci of harm.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Darüber hinaus rechnet die meisten Arbeit in diesem Bereich nicht mit der Intersektionalität, die die Idee beinhaltet, dass vielfältige soziale Identitäten Vorurteile kombinieren und einzigartige Opfer von Schaden sein können.", "score": 100.0}247{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_465.wav", "doc_id": "SUkmfOTvGi.seg_465", "src_text": "Let's get started.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "2003? Beginnen wir.", "score": 55.0}248{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_111.wav", "doc_id": "uZBWfYjYnf.seg_111", "src_text": "If we look at the main results of EDAtt, we'll plot the simultaneous speech translation results on graphs in which we have BLEU on one side that measures the translation quality, and average lagging that is the latency measure, and we also consider the computational aware average lagging that accounts for the model's computational times to predict the output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn wir uns die Hauptergebnisse ansehen, plotten wir die simultane Übersetzungsleistung in einem Diagramm, in dem wir auf einer Seite die Übersetzungsqualität blau und auf der anderen Seite die durchschnittliche Verzögerung d. h. die Latenzmessung messen. Wir berücksichtigen auch die computergesteuerte durchschnittliche Verzögerung. Daher möchten wir,", "score": 0.0}249{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_87.wav", "doc_id": "TVCREhgqUP.seg_87", "src_text": "We address this by inducing the alignment as part of the training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir beheben dies, indem wir die Ausrichtung als Teil der Ausbildung induzieren.", "score": 100.0}250{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_558.wav", "doc_id": "rISrKoXQCx.seg_558", "src_text": "So some preliminary results demonstrate that first, language models do have varying political leanings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Einige vorläufige Ergebnisse zeigen, dass erste Sprachmodelle unterschiedliche politische Bedeutungen haben.", "score": 97.0}251{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_324.wav", "doc_id": "dJGfOSFgZO.seg_324", "src_text": "For comparison, we also evaluated these conversations using three existing methods: Likert ratings on the turn-level, Likert ratings on the dialogue-level, and dialogue-level pairwise comparisons.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zur Vergleich haben wir diese Gespräche auch mit drei bestehenden Methoden bewertet: Lickert-Bewertungen auf der Drehungsebene, Lickert-Bewertungen auf der Dialogebene und Dialogebene: Paarweisen Vergleiche.", "score": 92.0}252{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_127.wav", "doc_id": "wLqFAuDnKa.seg_127", "src_text": "In this work, we present the first systematic study of large language model prompting for machine translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In dieser Arbeit stellen wir die erste systematische Untersuchung des Großsprachmodells für die maschinelle Übersetzung vor.", "score": 100.0}253{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_80.wav", "doc_id": "TVCREhgqUP.seg_80", "src_text": "To give you a teaser of the experimental results, here we compare our method with other treeless models on the COGS benchmark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Um Ihnen einen Eindruck von den experimentellen Ergebnissen zu vermitteln, vergleichen wir unsere Methode mit anderen Treelers-Modellen auf dem Korg-Benchmark.", "score": 60.0}254{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_270.wav", "doc_id": "PIZEXUFLAR.seg_270", "src_text": "However, there is no large-scale publicly-available multi-modal instruction task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "es ist kein groß angelegtes öffentlich verfügbares Multimodal", "score": 75.0}255{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_733.wav", "doc_id": "XejEJmgUmE.seg_733", "src_text": "So the minimal pair paradigm basically evaluates language models on top of acceptability judgments.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "So bewertet das minimale Paar-Paradigma grundlegend Sprachmodelle über Akzeptanzurteile, die", "score": 70.0}256{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_464.wav", "doc_id": "SUkmfOTvGi.seg_464", "src_text": "Today I'm going to present our paper Do CoNLL-2003 named entity taggers still work well in 2023?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "heute werde ich unseren Bericht präsentieren: Funktionieren die von Cornel 2003 benannten Entity-Tags noch gut im Jahr", "score": 60.0}257{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_315.wav", "doc_id": "dJGfOSFgZO.seg_315", "src_text": "Therefore, you might want to evaluate multiple dimensions of chat quality to understand the strengths and weaknesses of the model on a finer-grained level.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "daher könnten Sie mehrere Dimensionen der Dialogqualität bewerten, um die Stärken und Schwächen des Modells auf einem höheren Niveau zu verstehen.", "score": 95.0}258{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_390.wav", "doc_id": "WBLMIsdIrq.seg_390", "src_text": "So, depending on context, the meaning of the word changes, and therefore its translation changes as well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daher ändert sich die Bedeutung des Wortes je nach Kontext, und daher ändert sich auch seine Übersetzung.", "score": 100.0}259{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_746.wav", "doc_id": "XejEJmgUmE.seg_746", "src_text": "So we can do the same thing by choosing unacceptable sentences from the same matching, and that could also be used to test the models acceptability.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Also können wir dasselbe tun, indem wir inakzeptable Sätze aus der gleichen Übereinstimmung auswählen, und das könnte auch verwendet werden, um die Akzeptanz der Modelle zu testen.", "score": 80.0}260{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_257.wav", "doc_id": "oYCKgTzTDy.seg_257", "src_text": "To sum up, we build XSemPLR, a unified benchmark for cross-lingual semantic parsing with multiple natural languages and meaning representations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zusammenfassend bauen wir ein Beispiel, einen einheitlichen Benchmark für die Kreuzungssyntax mit mehreren natürlichen Sprachen und vielen Repräsentationen.", "score": 60.0}261{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_279.wav", "doc_id": "PIZEXUFLAR.seg_279", "src_text": "Ok, now I'm going to talk about multi-modal instruction tuning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Okay, ich werde über Multimodale Anweisungstuning sprechen.", "score": 65.0}262{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_53.wav", "doc_id": "TVCREhgqUP.seg_53", "src_text": "In this case, \"The girl slept.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "diesem Fall Übungen im Slip", "score": 50.0}263{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_497.wav", "doc_id": "SUkmfOTvGi.seg_497", "src_text": "We hope our paper calls for more research on how to improve generalizations of the models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir hoffen, dass unsere Arbeit weitere Forschungen darüber anregt, wie man die Verallgemeinerung der Modelle verbessern kann.", "score": 65.0}264{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_573.wav", "doc_id": "rISrKoXQCx.seg_573", "src_text": "And vice versa, right-leaning language models are better at detecting hate speech targeting white and men, however worse at detecting hate speech targeting at black LGBTQ plus and other minority communities.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und umgekehrt sind Rechtschreibmodelle besser, wenn es darum geht, weiße und männliche Sprache zu erkennen, aber schlechter, wenn es darum geht, schwarze, LGBTQ- und andere Minderheiten zu erkennen.", "score": 60.0}265{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_448.wav", "doc_id": "hgIDlKNiFM.seg_448", "src_text": "One based on the weight of CamemBERT and trained on a 4 GB set of NACHOS.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Einer basierte auf dem Gewicht von Camembert und trainierte auf vier Kilogramm von Natüron, ein anderer", "score": 60.0}266{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_274.wav", "doc_id": "PIZEXUFLAR.seg_274", "src_text": "For investigating multi-modal instruction tuning on our proposed dataset, we take OFA, a unified multi-modal pre-trained model, as our base model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Um die Multimodalsteuerung auf unserem vorgeschlagenen Datensatz zu untersuchen, nehmen wir OFA als unser Basismodell;", "score": 60.0}267{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_604.wav", "doc_id": "oeooqChmKK.seg_604", "src_text": "After a long day at work deciding cases in a law court, he was happy to relax.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "sie einen langen Tag mit dem Entscheiden von Fällen in einem Gerichtshof verbracht hatten. Er war froh, sich zu entspannen.", "score": 65.0}268{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_662.wav", "doc_id": "FLkGnzVRew.seg_662", "src_text": "We compare this to the other state-of-the-art AL strategies that are commonly used in the community.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir vergleichen dies mit anderen State-of-the-Art-Strategien, die in der Gemeinschaft allgemein verwendet werden. 'Nein, danke.", "score": 65.0}269{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_642.wav", "doc_id": "FLkGnzVRew.seg_642", "src_text": "High cognitive dissonance is also related to anxiety disorders and can help understand people's mental health better.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hohe Konstitutionsunterschiede hängen auch mit Angststörungen zusammen und können helfen, das mentale Wohlbefinden der Menschen besser zu verstehen.", "score": 60.0}270{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_44.wav", "doc_id": "aQpIWggfCo.seg_44", "src_text": "We hope the CoScript dataset can be a valuable resource to advance research on language planning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir hoffen, dass CoScript eine wertvolle Ressource sein kann, um die Forschung zur Sprachplanung voranzutreiben.", "score": 90.0}271{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_303.wav", "doc_id": "PIZEXUFLAR.seg_303", "src_text": "So overall, we propose the first large scale multi-model instruction tuning dataset with significantly improved their short capability of OFA, and we explore different transfer learning technique and show their benefits.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daher schlagen wir vor, ein erstes groß angelegtes Datenbanksystem für die Anpassung von Modellen zu erstellen, das die Fähigkeit der OFV erheblich verbessert und wir untersuchen verschiedene Techniken für die Übertragung von Lernfähigkeiten und zeigen ihre Vorteile.", "score": 50.0}272{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_786.wav", "doc_id": "WTTtiRKFZI.seg_786", "src_text": "OK.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "das", "score": 98.0}273{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_747.wav", "doc_id": "XejEJmgUmE.seg_747", "src_text": "And we can also do the same by choosing sentences from a different subset or a different data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und wir können dasselbe tun, indem wir Sätze aus einem anderen Untermenü oder Datensatz auswählen, also das,", "score": 93.0}274{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_220.wav", "doc_id": "oYCKgTzTDy.seg_220", "src_text": "Existing cross-lingual semantic parsing models are separately proposed and evaluated on data set of limited tasks and applications.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "so weiter. Bestehende Crosslingual-Semantik-Modellierungsmodelle werden separat vorgeschlagen und auf Datensätzen begrenzter Aufgaben und Anwendungen bewertet,", "score": 90.0}275{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_163.wav", "doc_id": "SLpqvupgvW.seg_163", "src_text": "Consider this alternative question.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Betrachten Sie diese alternative", "score": 70.0}276{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_470.wav", "doc_id": "SUkmfOTvGi.seg_470", "src_text": "At the same time, if we do observe poor generalization, what causes the performance drop of these models?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wenn wir eine schlechte Generalisierung feststellen, was verursacht den Leistungsabfall dieser Modelle?", "score": 100.0}277{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_830.wav", "doc_id": "GvEBWkLmuI.seg_830", "src_text": "In recent years, many have documented the prevalence of social bias and stereotypes in large language models, or LLMs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In den letzten Jahren haben viele die Prävalenzen von sozialen Vorurteilen und Stereotypen in großen Sprachmodellen dokumentiert.", "score": 90.0}278{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_769.wav", "doc_id": "XejEJmgUmE.seg_769", "src_text": "Please read our paper for more details of our experiments.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Bitte lesen Sie unser Papier für weitere Details zu unseren", "score": 97.0}279{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_630.wav", "doc_id": "oeooqChmKK.seg_630", "src_text": "If you're interested in more details, please see our paper and check out the data set and code on GitHub.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn Sie mehr Details interessieren, bitte sehen Sie sich unser Papier an und überprüfen Sie das Datensatz und den Code auf GitHub. ke", "score": 100.0}280{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_224.wav", "doc_id": "oYCKgTzTDy.seg_224", "src_text": "For example, there's only one single model to evaluate them.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Ibo, Yoruba, Afrikaans, Albanisch, Bosnisch, Serbisch, Kroatisch, Montenegrin, Maltesisch, Sardinisch, Katalanisch,", "score": 0.0}281{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_659.wav", "doc_id": "FLkGnzVRew.seg_659", "src_text": "\"Cumulative\" accumulates all the data collected from active annotation so far, whereas \"Iterative\" updates the model by training on the latest set of data collected.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daten, die aus aktiven Annotierungen gesammelt wurden, und Aktualisierung des Modells durch Training auf dem neuesten Satz von Daten,", "score": 74.0}282{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_372.wav", "doc_id": "gGbuDbHhyc.seg_372", "src_text": "To summarize, we showed that recent WSL approaches require clean, manually annotated samples for them to work properly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zusammenfassend zeigen wir, dass die neuesten WSL-Ansätze saubere, manuell annotierte Beispiele benötigen, um richtig zu funktionieren:", "score": 94.0}283{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_798.wav", "doc_id": "WTTtiRKFZI.seg_798", "src_text": "But it's also OK to say, \"Marge read yesterday this absolutely fascinating book about bees.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "über Bienen gelesen habe, ich sage, dass ich gestern March Read Yesterday, dieses absolut faszinierende Buch über", "score": 48.0}284{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_92.wav", "doc_id": "uZBWfYjYnf.seg_92", "src_text": "Hi, I'm Sara Papi from the University of Trento and Foundazione Bruno Kessler and I will briefly introduce the \"Attention as a Guide for Simultaneous Speech Translation\" paper, that is a joint work with Matteo Negri and Marco Turchi.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo, ich bin Sera Pappi von der Universität von Trento und der Stiftung Bruno Kessler, und ich werde kurz die Aufmerksamkeit als Leitfaden für das Simultanübersetzungspapier vorstellen, das eine gemeinsame Arbeit mit Matteo Negri und Marco Turchi ist.", "score": 92.0}285{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_339.wav", "doc_id": "dJGfOSFgZO.seg_339", "src_text": "And we look forward to seeing how conversational AI will advance in the coming months and years.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und wir freuen uns darauf, zu sehen, wie sich die kognitive KI in den kommenden Monaten und Jahren weiterentwickelt.", "score": 85.0}286{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_546.wav", "doc_id": "rISrKoXQCx.seg_546", "src_text": "Hi, I'm Shangbin, PhD student in the University of Washington.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "„Ich bin Doktorand an der Universität von Washington.", "score": 79.0}287{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_649.wav", "doc_id": "FLkGnzVRew.seg_649", "src_text": "On collecting around 1,000 examples of discourse unit pairs, we ran training for an initial classifier trained only on 43 examples of dissonance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "finden. Bei der Zusammenstellung von Tausenden von Beispielen für Diskussionseinheiten trainieren wir nur für die Klassifizierung", "score": 64.0}288{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_135.wav", "doc_id": "wLqFAuDnKa.seg_135", "src_text": "The difference observed is of more than one BLEURT points.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Differenz von SRF ist mehr als ein Blasenpunkt.", "score": 85.0}289{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_644.wav", "doc_id": "FLkGnzVRew.seg_644", "src_text": "Finally, cognitive dissonance is important to understand personal cognitive styles of individuals and helps us understand decision making processes better.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Kognitive Distanz ist wichtig, um persönliche kognitive Stile von Einzelpersonen zu verstehen, und hilft uns, Entscheidungsprozesse besser zu machen.", "score": 89.0}290{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_44.wav", "doc_id": "aQpIWggfCo.seg_44", "src_text": "We hope the CoScript dataset can be a valuable resource to advance research on language planning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir hoffen, dass die Scriptsammlung eine wertvolle Ressource für die weitere Forschung zur Sprachplanung sein kann.", "score": 89.0}291{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_614.wav", "doc_id": "oeooqChmKK.seg_614", "src_text": "Lastly, the \"Background-Inference\" setting, where both knowledge types are available only at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Drittens gibt es die Hintergrundausrichtung, bei der beide Wissensarten nur in der Interventionszeit verfügbar sind.", "score": 65.0}292{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_633.wav", "doc_id": "FLkGnzVRew.seg_633", "src_text": "I would like to present our work accepted into ACL 2023 as a long paper, \"Transfer Learning for Dissonance Detection: Addressing the Rare-Class Challenge.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ich würde gerne meine Arbeit als langes Papier mit dem Titel „Transfer Learning for Dissimilarity Detection“ vorstellen. Beginnend", "score": 72.0}293{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_216.wav", "doc_id": "oYCKgTzTDy.seg_216", "src_text": "Today I'm going to present our work \"XSemPLR: Cross-Lingual Semantic Parsing in Multiple Natural Languages and Meaning Representations\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Heute werde ich meine Arbeit vorstellen: Beispiel: Cross-Linguistic Semantic Parsing in mehreren natürlichen Sprachen und vielen Darstellungen. Das", "score": 79.0}294{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_786.wav", "doc_id": "WTTtiRKFZI.seg_786", "src_text": "OK.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "produzieren.", "score": 0.0}295{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_180.wav", "doc_id": "SLpqvupgvW.seg_180", "src_text": "Which is the alternative question.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das ist die alternative", "score": 53.0}296{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_528.wav", "doc_id": "dvGkKzmIaN.seg_528", "src_text": "The weight of the target embedding is proportional to the number of triggers in the sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das Gewicht des Ziel-Embeddings ist proportional zur Anzahl der Trigger im Satz.", "score": 83.0}297{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_397.wav", "doc_id": "WBLMIsdIrq.seg_397", "src_text": "To answer the first question, we started by measuring how much a word depends on context during translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Um die erste Frage zu beantworten, beginnen wir damit, zu messen, wie stark ein Wort von dem Kontext während der Übersetzung abhängt.", "score": 100.0}298{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_159.wav", "doc_id": "SLpqvupgvW.seg_159", "src_text": "Hi!", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hi,", "score": 100.0}299{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_711.wav", "doc_id": "oaOHnMCwad.seg_711", "src_text": "We find that Dynahate is also most aligned to English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir stellen fest, dass Dianet Heat ebenfalls am meisten mit englischsprachigen Ländern übereinstimmt.", "score": 90.0}300{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_628.wav", "doc_id": "oeooqChmKK.seg_628", "src_text": "However, with task-specific training, some models successfully integrate knowledge from multiple sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Allerdings können einige Modelle mit task-spezifischer Ausbildung erfolgreich Wissen aus mehreren Quellen integrieren.", "score": 100.0}301{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_351.wav", "doc_id": "gGbuDbHhyc.seg_351", "src_text": "Technically, this claim is not wrong, but there's a catch, which is that people do assume that there's an additional clean validation set available for model selection.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Technisch gesehen ist dieser Anspruch nicht falsch, aber es gibt einen Haken. Das heißt, dass die Menschen davon ausgehen, dass es eine zusätzliche saubere Validierungsmethode für die Modellauswahl gibt. Wir werfen", "score": 83.0}302{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_343.wav", "doc_id": "gGbuDbHhyc.seg_343", "src_text": "This is joint work with Xiaoyu Shen, Marius Mosbach, Andreas Stephan, and Dietrich Klakow.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Lernfähigkeit. Das ist eine gemeinsame Arbeit mit Shaul Usch, Mario Muzspara, Andreas Stefan und Dietrich Klakow.", "score": 64.0}303{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_494.wav", "doc_id": "SUkmfOTvGi.seg_494", "src_text": "At the same time, we also found that the performance drop here is caused by temporal drift and kind of surprisingly, it is not caused by adaptive overfitting even though CoNLL-2003 has been used for over 20 years.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Gleichzeitig haben wir auch festgestellt, dass die Leistungsabnahme hier durch zeitliche Drift verursacht wird und, was ziemlich überraschend ist, nicht durch adaptives Überpassen. Obwohl \"Corne 2003\" seit mehr als zwanzig Jahren verwendet", "score": 85.0}304{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_330.wav", "doc_id": "dJGfOSFgZO.seg_330", "src_text": "You can see how the combination of all ABC-Eval metrics explains over 25% of conversation quality, and as you remove the metrics one at a time, most of them result in losing a decent amount of information about the quality.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Sie können sehen, wie die Kombination aller ABC-EVA-Metriken über 25 Prozent der Gesprächsqualität erklärt und wenn Sie die Metriken einzeln entfernen, verlieren die meisten von ihnen einen guten Teil der Informationen über die Qualität. Auf", "score": 100.0}305{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_731.wav", "doc_id": "XejEJmgUmE.seg_731", "src_text": "This is a joint work with John Gauthier, Aaron Mueller, Kanishka Misra, Karen Fences, Roger Levy, and Adina Williams.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "dürfen. Es ist eine gemeinsame Arbeit mit John Gauthier, Aaron Mueller, Kaniška Mistrá, Káren Fuentés, Roger Levy und Adina Williams.", "score": 100.0}306{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_571.wav", "doc_id": "rISrKoXQCx.seg_571", "src_text": "So we see that if we investigate the per category performance, that is to say if we separate the performance into different demographics or political leaning of news media we can see a pattern.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Leistungskategorie untersuchen, das heißt, wenn wir die Leistung trennen. Unterschiedliche Demografien oder politische Nachrichtenmedien können zeigen, dass beispielsweise Sprachdetektionsmodelle bessere", "score": 92.0}307{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_447.wav", "doc_id": "hgIDlKNiFM.seg_447", "src_text": "In addition to this comparison, we introduced three models trained on continual pre-training to analyze the impact of pre-training strategy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zusätzlich zu diesem Vergleich führen wir drei Modelle des kontinuierlichen Trainings ein, um die Auswirkungen der Trainingsstrategie zu analysieren.", "score": 86.0}308{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_628.wav", "doc_id": "oeooqChmKK.seg_628", "src_text": "However, with task-specific training, some models successfully integrate knowledge from multiple sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Mit einer taskspezifischen Schulung integrieren jedoch einige Modelle erfolgreich Wissen aus mehreren Quellen. Trotzdem", "score": 100.0}309{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_487.wav", "doc_id": "SUkmfOTvGi.seg_487", "src_text": "For data overfitting, we saw that from the graph on the right, the red best fit line has a gradient that is greater than one.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Bei der Anpassung der Überpassung haben wir festgestellt, dass die rote Bestpassungslinie auf der rechten Seite eine Steigung von mehr als eins hat.", "score": 65.0}310{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_218.wav", "doc_id": "oYCKgTzTDy.seg_218", "src_text": "And Cross-Lingual Semantic Parsing is the task to translate queries in multiple natural languages into multiple meaning representations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Aufgabe besteht darin, Aussagen in mehreren natürlichen Sprachen in mehrere Bedeutungsrepräsentationen zu übersetzen.", "score": 95.0}311{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_787.wav", "doc_id": "WTTtiRKFZI.seg_787", "src_text": "The argument is based on the principle of dependency length minimization that I will explain on the basis of these examples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Ok, das Argument stützt sich auf das Prinzip der Abhängigkeitslängenminimierung, das wir auf der Grundlage dieser Beispiele erklären. So", "score": 70.0}312{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_310.wav", "doc_id": "dJGfOSFgZO.seg_310", "src_text": "And today we'll tell you all about ABC-Eval, a new dimensional approach to evaluating conversational AI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und heute werden wir Ihnen alles über ABC-Eval erzählen, einen neuen dimensionalen Ansatz zur Bewertung der konversationalen AI.", "score": 100.0}313{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_41.wav", "doc_id": "aQpIWggfCo.seg_41", "src_text": "In summary, we establish the constrained language planning problem.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zusammenfassend können wir sagen, dass wir das Problem der begrenzten Sprachplanung aufgestellt", "score": 100.0}314{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_507.wav", "doc_id": "dvGkKzmIaN.seg_507", "src_text": "For example, OpenAI offers a GPT based embedding API.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zum Beispiel bietet OpenLayers eine GPX-basierte Einbettungs-API. Jedoch haben jüngste", "score": 89.0}315{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_783.wav", "doc_id": "WTTtiRKFZI.seg_783", "src_text": "So we get dependencies from the governor.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "bekommen wir Abhängigkeiten von dem", "score": 60.0}316{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_124.wav", "doc_id": "wLqFAuDnKa.seg_124", "src_text": "PaLM is a 540 billion-parameter large language model presented last year in 2022.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "FARM ist ein 500-Milliarden-Parameter-Modell der großen Sprache, das im Jahr 2022 vorgestellt wurde.", "score": 55.0}317{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_134.wav", "doc_id": "wLqFAuDnKa.seg_134", "src_text": "The majority of sentences 516 out of 1,000.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die Mehrheit der Sätze, sechzehn von tausend,", "score": 88.0}318{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_191.wav", "doc_id": "SLpqvupgvW.seg_191", "src_text": "The second one is when the entities have similar titles, for example, two books with the name \"The Return\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die zweite ist, wenn die Entitäten ähnliche Titel haben, zum Beispiel zwei Bücher mit dem Namen „Der Retter“", "score": 95.0}319{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_34.wav", "doc_id": "aQpIWggfCo.seg_34", "src_text": "We appy our method for building a dataset of constrained language planning, named as CoScript.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir planen unsere Methode für den Aufbau eines Datensatzes für die kontrollierte Sprachplanung, genannt „Codescript“.", "score": 88.0}320{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_613.wav", "doc_id": "oeooqChmKK.seg_613", "src_text": "Second, there's a \"Background-Both\" setting, where background knowledge is available both at pretrain time and inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zweitens gibt es eine Hintergrundbeleuchtung, wobei die Hintergrundwissen sowohl zu Trainingszeiten als auch zu Interferenzzeiten verfügbar sind;", "score": 100.0}321{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_429.wav", "doc_id": "WBLMIsdIrq.seg_429", "src_text": "Thank you so much for your attention.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Vielen Dank für Ihre Aufmerksamkeit,", "score": 100.0}322{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_365.wav", "doc_id": "gGbuDbHhyc.seg_365", "src_text": "But that's not the end of the story, because if we either way decide to access clean samples, then training on them directly will even achieve better performance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Aber das ist nicht das Ende der Geschichte, denn wenn wir uns entscheiden, saubere Proben zu verwenden, wird das Training damit sogar noch bessere Ergebnisse erzielen.", "score": 96.0}323{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_801.wav", "doc_id": "WTTtiRKFZI.seg_801", "src_text": "So here we have a dependency from \"read\" to the adjunct of length 7 measured in words and from \"read\" to \"book\" of length 4, so together it's 11.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "haben wir eine Abhängigkeit von red zu dem Ende von length sieben gemessen in Wörtern und von red zu book von length vier, also zusammen elf, wenn", "score": 94.0}324{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_600.wav", "doc_id": "oeooqChmKK.seg_600", "src_text": "Here is an example from our data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "ein Beispiel aus unserem Datensatz:", "score": 98.0}325{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_37.wav", "doc_id": "aQpIWggfCo.seg_37", "src_text": "This figure shows the constraint distribution of CoScript.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Diese Abbildung zeigt eine eingeschränkte Verteilung von CoScript.", "score": 100.0}326{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_384.wav", "doc_id": "WBLMIsdIrq.seg_384", "src_text": "A Data-driven, Multilingual Exploration\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "- a data-driven multilingual exploration' vorstellen.", "score": 95.0}327{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_19.wav", "doc_id": "aQpIWggfCo.seg_19", "src_text": "The heat map in the figure shows that the planning performance of InstructGPTs varies considerably for goals of different categories.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "die Karteikarte zeigt, dass die Planungsleistung von Unterrichtseinrichtungen für Mädchen unterschiedlicher Kategorien beträchtlich variiert.", "score": 20.0}328{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_204.wav", "doc_id": "SLpqvupgvW.seg_204", "src_text": "For example, \"the one without words\", \"not the one with the 12 year old boy\", or \"the fictional one\", or \"comes from Azerbaijan\", and so on.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "zum Beispiel die ohne Worte, nicht der mit dem zwölfjährigen Jungen, oder die fiktive, die aus Aserbaidschan kommt und so weiter.", "score": 75.0}329{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_425.wav", "doc_id": "WBLMIsdIrq.seg_425", "src_text": "But these models are not much better than models that do not use context on other phenomena like ellipsis, pronouns, and verb form.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "diese Modelle sind nicht viel besser als die Modelle, die keine Kontexte auf anderen", "score": 94.0}330{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_280.wav", "doc_id": "PIZEXUFLAR.seg_280", "src_text": "So for the training dataset, we use 53 tasks from 9 groups for training and we sample 10,000 instances per task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Für das Trainingsdatenset verwenden wir fünfunddreißig Aufgaben aus der N-Gruppe für das Training und beispielhaft zehntausend Instanzen", "score": 30.0}331{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_283.wav", "doc_id": "PIZEXUFLAR.seg_283", "src_text": "In addition, we randomly sample 20 tasks from the test split of natural instructions as an unseen task for NLP.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "jede Aufgabe; zudem nehmen wir zufällig eine Aufgabe aus dem Testset der natürlichen Anweisung als unsichtbare Aufgabe für den NLP.", "score": 78.0}332{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_528.wav", "doc_id": "dvGkKzmIaN.seg_528", "src_text": "The weight of the target embedding is proportional to the number of triggers in the sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das Gewicht des Ziel-Embeddings ist proportional zur Anzahl der Trigger im Satz.", "score": 90.0}333{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_866.wav", "doc_id": "GvEBWkLmuI.seg_866", "src_text": "So for example, the words describing Latina women include things like \"vibrant\" and \"curvaceous\" which connect to a trope of tropicalism.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "der Farbe, wie zum Beispiel die lateinische Frau, die lebendig und lebhaft ist.", "score": 40.0}334{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_222.wav", "doc_id": "oYCKgTzTDy.seg_222", "src_text": "But Chinese is missing and lack of coverage on certain meaning representation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Usbekisch, Kirgisisch, Mongolisch, Tibetisch, Nepali, Bengali, Marathi, Gujarati, Telugu, Tamil,", "score": 0.0}335{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_490.wav", "doc_id": "SUkmfOTvGi.seg_490", "src_text": "So what about temporal drift then?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Was ist mit der Temperaturdiffusion? Für", "score": 10.0}336{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_713.wav", "doc_id": "oaOHnMCwad.seg_713", "src_text": "So for GPT 4, in the social acceptability task, we find that it's most aligned to people with a college education or Graduate School education and we find the same for Dynahate where it's most aligned to people with a college education.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "eine Hochschulausbildung haben, so dass für die Aufgabe der sozialen Eingängigkeit vier GBD finden, dass es die meisten Verbindungen mit Personen mit einer Hochschulausbildung oder einer Abschluss-Hochschulausbildung gibt. Und wir finden das Gleiche für Donny Haide, wo es den Menschen mit einer Hochschulbildung am meisten zusagt.", "score": 10.0}337{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_647.wav", "doc_id": "FLkGnzVRew.seg_647", "src_text": "Tweets were passed using the PDTB parser, and pairs of discourse units were annotated according to the guidelines that are described in our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die Tweets wurden unter Verwendung eines PTB-Parsers und Paare von Diskurs-Einheiten gemäß den Richtlinien, die im Papier beschrieben sind, analysiert.", "score": 100.0}338{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_297.wav", "doc_id": "PIZEXUFLAR.seg_297", "src_text": "So we also did one experiment.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "So führen wir auch ein Experiment durch,", "score": 90.0}339{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_119.wav", "doc_id": "uZBWfYjYnf.seg_119", "src_text": "If you want to discover more results, read our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wenn Sie mehr Ergebnisse entdecken möchten, lesen Sie unser Papier,", "score": 100.0}340{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_710.wav", "doc_id": "oaOHnMCwad.seg_710", "src_text": "So for the GPT 4 social acceptability analysis, we find that it's most aligned to confucian and English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "dass für die Analyse der sozialen Akzeptabilität der GPD 4 festgestellt wird, dass es am meisten mit Konflikt und englischsprachigen Ländern übereinstimmt, und", "score": 60.0}341{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_804.wav", "doc_id": "WTTtiRKFZI.seg_804", "src_text": "That's why this sounds quite okay.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "klingt das in Ordnung.", "score": 100.0}342{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_630.wav", "doc_id": "oeooqChmKK.seg_630", "src_text": "If you're interested in more details, please see our paper and check out the data set and code on GitHub.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wenn Sie mehr Einzelheiten erfahren möchten, sehen Sie sich bitte unser Paper und den Datensatz auf GitHub an.", "score": 80.0}343{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_870.wav", "doc_id": "GvEBWkLmuI.seg_870", "src_text": "And while it sounds positive at first glance, there's been work showing that this kind of archetype actually is very harmful because it puts a lot of pressure on these demographics to be resilient and strong against societal obstacles.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "und klingt auf den ersten Blick positiv. Es hat sich gezeigt, dass diese Art von Archetypus eigentlich sehr schädlich ist, weil sie viel Druck auf diese Demografien ausübt, um widerstandsfähig und stark gegen soziale Hindernisse zu sein.", "score": 95.0}344{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_95.wav", "doc_id": "uZBWfYjYnf.seg_95", "src_text": "And what are the problems of the current SimulST models?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und was sind die Probleme der aktuellen SimulS-Modelle?", "score": 100.0}345{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_588.wav", "doc_id": "rISrKoXQCx.seg_588", "src_text": "Thank you for your time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dank für Ihre Zeit.", "score": 100.0}346{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_693.wav", "doc_id": "oaOHnMCwad.seg_693", "src_text": "Our framework works in two main steps.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Unser Rahmenwerk funktioniert in zwei Hauptstufen.", "score": 100.0}347{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_703.wav", "doc_id": "oaOHnMCwad.seg_703", "src_text": "We've then compared these, annotations with Social Chemistry, Delphi and GPT 4.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir vergleichen dann diese Anmerkungen mit Social Chemistry, Delphi und GPT-4. Wir wiederholen", "score": 90.0}348{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_491.wav", "doc_id": "SUkmfOTvGi.seg_491", "src_text": "For temporal drift, we did an experiment to retrain or continue to pre-train some models with more recent data and we found that the performance degrades with larger temporal gap and this confirms our hypothesis that the main cause of the performance drop is temporal drift.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "die zeitliche Drift haben wir ein Experiment durchgeführt, um einige Modelle mit neueren Daten neu zu trainieren oder weiter zu pre-trainieren, und wir haben festgestellt, dass die Leistung mit größeren zeitlichen Lücken abnimmt. Dies bestätigt unsere Hypothese, dass die Hauptursache für den Leistungsabfall die zeitliche Drift ist.", "score": 100.0}349{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_175.wav", "doc_id": "SLpqvupgvW.seg_175", "src_text": "Our data set collection methodology emphasizes informality using a cartoon completion setup.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unsere Datensatzsammlungsmethode betont die Informalität mit einem Cartoon-Completion-Setup.", "score": 95.0}350{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_764.wav", "doc_id": "XejEJmgUmE.seg_764", "src_text": "And after doing like several of these perturbations, we find that none of these noises are actually making the model like change its course in terms of how it shows us the MPP judgement print.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "„Lärm“ zu versehen, und nachdem wir viele dieser Störungen durchgeführt hatten Wir stellen fest, dass keiner dieser Geräusche tatsächlich den Modellverlauf in Bezug auf die Art und Weise, wie es", "score": 59.0}351{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_213.wav", "doc_id": "SLpqvupgvW.seg_213", "src_text": "Here is a link to our dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hier ist ein Link zu unseren Datensätzen.", "score": 96.0}352{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_290.wav", "doc_id": "PIZEXUFLAR.seg_290", "src_text": "If it's a multi-modal generation task, we report Rouge-L. For NLP task, we report Rouge-L as well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "es sich um eine multimodale Generierungsaufgabe handelt, berichten wir über RUGL, und für NPR-Aufgaben berichten wir auch über RUGL.", "score": 60.0}353{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_392.wav", "doc_id": "WBLMIsdIrq.seg_392", "src_text": "Firstly because only a small portion of translations depend on context which makes corpus-level metrics like BLEU unable to capture these translations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist jedoch ziemlich schwer: Erstens, weil nur ein kleiner Teil der Übersetzungen vom Kontext abhängt, was Korpus-Ebenen-Metriken wie Blue daran hindert, diese Übersetzungen zu erfassen. Und", "score": 91.0}354{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_423.wav", "doc_id": "WBLMIsdIrq.seg_423", "src_text": "This again demonstrates that it is difficult to determine the best document-level translation system if we use corpus-level metrics alone.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies zeigt, dass es schwierig ist, das beste Dokumentenverarbeitungssystem zu ermitteln, wenn man die Korpus-Metrik verwendet. Jetzt", "score": 60.0}355{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_723.wav", "doc_id": "oaOHnMCwad.seg_723", "src_text": "I mean, we want to emphasise that inclusive NLP isn't just making.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "ist die Maska-Initiative. Ich möchte betonen, dass inklusives P nicht nur für alle", "score": 61.0}356{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_811.wav", "doc_id": "WTTtiRKFZI.seg_811", "src_text": "So when the difference between the lengths of the two conjuncts grows, the shorter conjunct prefers to be the first one, stronger, right?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "also der Unterschied zwischen den Längen der beiden Konjunktionen groß ist, dann ist der kürzere Konjunkt die erste stärkere, also ist", "score": 60.0}357{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_193.wav", "doc_id": "SLpqvupgvW.seg_193", "src_text": "And finally when they have similar info boxes or attributes on Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und schließlich, wenn sie ähnliche Infoboxen oder Attribute auf Wikipedia haben,", "score": 100.0}358{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_142.wav", "doc_id": "wLqFAuDnKa.seg_142", "src_text": "And when we go, as in our case, to five-shot prompting, there is nearly no difference to the actual form of the prompting.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wenn wir in unserem Fall zum Schießen gehen, gibt es fast keine Unterschiede zur tatsächlichen Form des Schießens. Es", "score": 29.0}359{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_802.wav", "doc_id": "WTTtiRKFZI.seg_802", "src_text": "When you swap these two constituents, the sum of these two dependencies becomes 6.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Sie diese beiden Konstituenten verschieben, dann werden einige dieser beiden Abhängigkeiten zu", "score": 58.0}360{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_331.wav", "doc_id": "dJGfOSFgZO.seg_331", "src_text": "On the other hand, the combination of all turn-level Likert metrics explains far less of the quality, and fewer of these metrics carry unique information.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "der anderen Seite erklärt die Kombination alternativer Lickert-Metriken viel weniger der Qualität und weniger dieser Metriken tragen einzigartige Informationen.", "score": 93.0}361{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_849.wav", "doc_id": "GvEBWkLmuI.seg_849", "src_text": "So for instance, the word \"warrior\" is usually associated with men.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "So ist das Wort „Mann“ oder „Krieger“ normalerweise mit „Mann“ verbunden,", "score": 93.0}362{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_598.wav", "doc_id": "oeooqChmKK.seg_598", "src_text": "We introduce a coreference resolution task, designed to probe for the ability to draw on knowledge available in different sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir führen eine Korrelationsanalyse durch, die darauf ausgelegt ist, die Fähigkeit zu bewerten, auf Wissen in verschiedenen Quellen zuzugreifen.", "score": 85.0}363{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_865.wav", "doc_id": "GvEBWkLmuI.seg_865", "src_text": "Furthermore, there's a lot of common tropes that are reflected in these words, especially for women of color.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und noch mehr sind die vielen Gemeinsamkeiten, die in diesen Wörtern enthalten sind, insbesondere für eine", "score": 63.0}364{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_494.wav", "doc_id": "SUkmfOTvGi.seg_494", "src_text": "At the same time, we also found that the performance drop here is caused by temporal drift and kind of surprisingly, it is not caused by adaptive overfitting even though CoNLL-2003 has been used for over 20 years.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Gleichzeitig stellten wir auch fest, dass der Leistungsverlust hier durch zeitliche Schwankungen verursacht wird, und überraschenderweise ist er nicht durch adaptiven Überdrehungskoeffizienten verursacht, obwohl der Cornal 2003 seit über zwanzig Jahren verwendet wird.", "score": 64.0}365{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_564.wav", "doc_id": "rISrKoXQCx.seg_564", "src_text": "For example, for RoBERTa further trained on the left-leaning Reddit corpus we can see a substantial liberal shift in terms of its political biases.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zum Beispiel können wir für Roberta, die eine weiter gefällige Stimme hat und weiter auf dem linken Korpus trainiert ist, einen wesentlichen liberalen Wandel in ihrer Stimme sehen. In Bezug auf seine politischen Überzeugungen.", "score": 64.0}366{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_294.wav", "doc_id": "PIZEXUFLAR.seg_294", "src_text": "As we can see, instruction tuning can significantly improve OFA's performance on seen multi-modal tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wie wir sehen können, dass die Anpassung der Anweisungen die Leistung von OIS auf ähnlichen Multimodal-Aufgaben erheblich verbessern kann.", "score": 90.0}367{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_779.wav", "doc_id": "WTTtiRKFZI.seg_779", "src_text": "Now those are asymmetric approaches to coordinate structures, such as the Prague approach.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "symmetrischen Ansätze für Koordinatenstrukturen, wie z. B. der", "score": 65.0}368{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_405.wav", "doc_id": "WBLMIsdIrq.seg_405", "src_text": "First, we look at part-of-speech tags that have high mean P-CXMI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zunächst sehen wir uns die Sprachetiketten an, die hohe Minen haben, wie z.B.", "score": 64.0}369{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_165.wav", "doc_id": "SLpqvupgvW.seg_165", "src_text": "Here, a user wants to select between one of these two songs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hier möchte ein Benutzer zwischen zwei dieser beiden Lieder wählen.", "score": 100.0}370{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_418.wav", "doc_id": "WBLMIsdIrq.seg_418", "src_text": "We then use the MuDA tagger, by applying the tagger on a parallel corpus that we want to use for evaluation and we apply our translation metrics of choice on the context-dependent examples that the MuDA tagger has identified.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "haben. Dann verwenden wir den Muta-Tagger, indem wir den Tagger auf das parallele Korpus anwenden, das wir für die Bewertung verwenden möchten. Und wir wenden unsere Übersetzungsmetriken zur Auswahl auf die kontextabhängigen Beispiele an, die der Modus-Tagger identifiziert hat.", "score": 73.0}371{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_77.wav", "doc_id": "TVCREhgqUP.seg_77", "src_text": "Then we jump to the next multiset token, to determine the second token in the output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dann springen wir zum nächsten Multiset-Token, um den zweiten Token im Output zu bestimmen.", "score": 95.0}372{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_777.wav", "doc_id": "WTTtiRKFZI.seg_777", "src_text": "Right.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "die", "score": 0.0}373{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_122.wav", "doc_id": "wLqFAuDnKa.seg_122", "src_text": "Hello everyone, my name is David Vilar, and I will be giving a short review of the paper \"Prompting PaLM for Translation: Assessing Strategies and Performance.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo, mein Name ist Said, und ich werde eine kurze Übersicht über das Papier „Translation, Assessing Strategies and Performance“ geben.", "score": 68.0}374{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_869.wav", "doc_id": "GvEBWkLmuI.seg_869", "src_text": "This connects to an archetype that people have called the \"Strong Black Women\" archetype.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dies verbindet sich mit einem Archetyp, den die Menschen als den starken schwarzen Archetyp bezeichnet haben,", "score": 83.0}375{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_113.wav", "doc_id": "uZBWfYjYnf.seg_113", "src_text": "But also we want that they are shifted on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Aber auch wir wollen, dass sie auf die linke Seite versetzt werden", "score": 100.0}376{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_755.wav", "doc_id": "XejEJmgUmE.seg_755", "src_text": "We increase the context length toward up to 1024 for to max out OPT and GPT 2 models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir verlängerten den Kontext auf bis zu 10.000, um die OPT- und GPT-2-Modelle zu maximieren, und", "score": 76.0}377{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_360.wav", "doc_id": "gGbuDbHhyc.seg_360", "src_text": "Otherwise, there is a large performance drop.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ansonsten gibt es einen großen Leistungsabfall,", "score": 100.0}378{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_327.wav", "doc_id": "dJGfOSFgZO.seg_327", "src_text": "In addition, ABC-Eval labels are more predictive of the overall conversation quality compared to metrics produced by existing methods, as shown by this simple linear regression analysis.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Darüber hinaus sind ABC-EV-Labels im Hinblick auf die Gesprächsqualität der Gesamtkommunikation vorhersagbarer als Metriken, die von existierenden Methoden erzeugt werden, wie es durch diese einfachen linearen Regressionen gezeigt", "score": 100.0}379{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_352.wav", "doc_id": "gGbuDbHhyc.seg_352", "src_text": "We can't stop on this problem setting, but this implies that additional manual annotations are required in weakly supervised learning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Validierungsset gibt. Wir haben Zweifel an dieser Problemstellung, aber das impliziert, dass zusätzliche manuelle Anmerkungen beim Erlernen von Wikis erforderlich sind,", "score": 68.0}380{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_167.wav", "doc_id": "SLpqvupgvW.seg_167", "src_text": "But sometimes an indirect reference is more appropriate to have a more natural conversation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "erste. Aber manchmal ist ein indirekter Hinweis angemessener, um eine natürlichere Unterhaltung zu haben;", "score": 92.0}381{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_316.wav", "doc_id": "dJGfOSFgZO.seg_316", "src_text": "One approach is to simply ask human judges to evaluate several dimensions of dialogue quality, such as the relevance of model responses using existing comparative or Likert scale methods.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Eine Möglichkeit besteht darin, einfach Menschen zu bitten, mehrere Dimensionen der Dialogqualität zu bewerten, wie die Relevanz der Modellantworten, mithilfe bestehender vergleichender oder Likert-Skala-Methoden.", "score": 97.0}382{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_132.wav", "doc_id": "wLqFAuDnKa.seg_132", "src_text": "Finally, we provide some recommendations for prompt selection strategies.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "geben wir einige Empfehlungen für prompte Auswahlstrategien. Die Stimulierung", "score": 71.0}383{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_844.wav", "doc_id": "GvEBWkLmuI.seg_844", "src_text": "Our prompts to generate these personas were inspired by a study where they gave these prompts to human subjects, finding that by giving it to human subjects, they also were able to surface racial stereotypes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Unsere Proben, die diese Persönlichkeiten hervorbrachten, wurden von einer Studie inspiriert, in der diese Proben an menschlichen Subjekten getestet wurden, wobei festgestellt wurde, dass sie auch rassenspezifische Stereotypen aufweisen. Und", "score": 86.0}384{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_736.wav", "doc_id": "XejEJmgUmE.seg_736", "src_text": "And then the hope is that the model, basically, puts more probability to the acceptable sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "dann hofft man, dass das Modell im Wesentlichen mehr Wahrscheinlichkeit für den akzeptablen Satz hat. Die derzeitige Pipeline", "score": 79.0}385{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_757.wav", "doc_id": "XejEJmgUmE.seg_757", "src_text": "Now, what happens when we choose sentences from the same data set?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Was passiert nun, wenn wir Sätze aus dem gleichen Datensatz auswählen?", "score": 100.0}386{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_145.wav", "doc_id": "wLqFAuDnKa.seg_145", "src_text": "So it's important to select the examples from high-quality translations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist also wichtig, Beispiele aus hochwertigen Übersetzungen auszuwählen,", "score": 100.0}387{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_304.wav", "doc_id": "PIZEXUFLAR.seg_304", "src_text": "We design a new metric called sensitivity.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir entwerfen ein neues Maß, das Sensitivity genannt wird.", "score": 62.0}388{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_518.wav", "doc_id": "dvGkKzmIaN.seg_518", "src_text": "Therefore, in this paper we propose Embedding marker, which is a backdoor based watermark method applicable to embedding as services.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "schlagen wir in dieser Arbeit eine Implementierungsmarker vor, die eine Backdoor-basierte Wasserzeichenmethode ist, die auf die Implementierung von ADS-Diensten anwendbar ist.", "score": 80.0}389{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_589.wav", "doc_id": "oeooqChmKK.seg_589", "src_text": "Hello everyone, I'm Akshatha, and today my co-author Martin and I are presenting our work \"The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo alle, ich bin Amratha Mathew und ich präsentiere heute meine Arbeit, die KIT Masterclass: Evaluierung der Wissensintegration aus mehreren Quellen.", "score": 30.0}390{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_9.wav", "doc_id": "aQpIWggfCo.seg_9", "src_text": "A good planner should write scripts that are reasonable and faithful to constraints.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ein guter Planer sollte Skripte schreiben, die vernünftig und den Einschränkungen treu sind.", "score": 99.0}391{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_635.wav", "doc_id": "FLkGnzVRew.seg_635", "src_text": "Simply put, cognitive dissonance is two beliefs or actions that are inconsistent, such as this example where a person states, \"I know that cigarettes could kill me\", and then goes on to say \"I grabbed a couple of smokes after the meeting\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "studieren, ist kognitive Dissonanz einfach zwei Überzeugungen oder Handlungen, die inkonsistent sind. Dieses Beispiel zeigt, dass ich weiß, dass die Zigaretten mich umbringen würden, und dann sage ich, dass ich nach dem Treffen ein paar Rauchpausen einlegen würde.", "score": 75.0}392{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_105.wav", "doc_id": "uZBWfYjYnf.seg_105", "src_text": "Our solution is to propose EDAtt, or Encoder-Decoder Attention, and it is a strategy for which we decide whether to emit or not a partial translation, based on where attention points to.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unsere Lösung besteht darin, einen EDA- oder einen Encoder für die Codierung der Aufmerksamkeit zu vorschlagen, und es ist eine Strategie, bei der wir entscheiden, ob wir eine partielle Übersetzung senden oder nicht, basierend auf der Position der Aufmerksamkeit.", "score": 50.0}393{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_756.wav", "doc_id": "XejEJmgUmE.seg_756", "src_text": "And we saw here in the orange dotted line, the MPP judgments are relatively stable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wir sahen hier in der orangefarbenen Zeile, dass die MPP-Urteile relativ stabil sind.", "score": 100.0}394{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_746.wav", "doc_id": "XejEJmgUmE.seg_746", "src_text": "So we can do the same thing by choosing unacceptable sentences from the same matching, and that could also be used to test the models acceptability.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "hinzufügen. Daher können wir dasselbe tun, indem wir unannehmbare Sätze aus demselben Matching auswählen, und das könnte auch verwendet werden, um die Akzeptanzfähigkeit der Modelle zu testen.", "score": 89.0}395{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_804.wav", "doc_id": "WTTtiRKFZI.seg_804", "src_text": "That's why this sounds quite okay.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "klingt das ganz in", "score": 8.0}396{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_312.wav", "doc_id": "dJGfOSFgZO.seg_312", "src_text": "So let's say that you just developed a dialogue model and you want to see how well it compares against the current state-of-the-art.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Sie uns also sagen, dass Sie gerade ein Dialogmodell entwickelt haben und sehen möchten, wie gut es mit dem aktuellen Stand der Technik vergleichbar", "score": 98.0}397{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_626.wav", "doc_id": "oeooqChmKK.seg_626", "src_text": "Additional experiments with fictional knowledge indicated even the best performing models, cannot reliably integrate backward knowledge provided only at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zusätzliche Experimente mit fiktivem Wissen zeigen, dass selbst die besten Modelle das Hintergrundwissen, das ihnen zur Verfügung gestellt wird, nicht zuverlässig integrieren können.", "score": 75.0}398{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_193.wav", "doc_id": "SLpqvupgvW.seg_193", "src_text": "And finally when they have similar info boxes or attributes on Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und schließlich, wenn sie ähnliche Infoboxen oder Attribute auf Wikipedia haben,", "score": 100.0}399{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_717.wav", "doc_id": "oaOHnMCwad.seg_717", "src_text": "So, given that there is positionality in NLP, what can we do about it?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Heat Task Analysen. Also, was können wir tun, wenn", "score": 56.0}400{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_292.wav", "doc_id": "PIZEXUFLAR.seg_292", "src_text": "So this measures the model's ability to consistently produce the same outputs for the same task regardless of the slight variation in the wording of the instruction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Fähigkeit des Modells misst, unabhängig von geringfügigen Abweichungen in der Anweisung immer die gleichen Ergebnisse für die gleiche Aufgabe zu liefern. Hier sind unsere Hauptergebnisse, wie man sieht, kann die", "score": 100.0}401{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_675.wav", "doc_id": "oaOHnMCwad.seg_675", "src_text": "I'm Jenny, a first year PhD student at Carnegie Mellon University and today I'll be presenting your work NLPositionality characterising design biases of datasets and Models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Jenny, Studentin im ersten Jahr an der Universität Karnegi-Mellon, und werde heute meine Arbeit und ihre Position beschreiben, indem ich Datenmodelle erstelle.", "score": 45.0}402{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_543.wav", "doc_id": "dvGkKzmIaN.seg_543", "src_text": "That's all.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das ist alles, danke.", "score": 95.0}403{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_624.wav", "doc_id": "oeooqChmKK.seg_624", "src_text": "When trained on KITMUS, however, both C2F and BERT4Coref perform significantly better than the random choice.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "nicht gut. Das Training auf dem Kidmus hingegen ist bei beiden Modellen signifikant besser als die zufällige Wahl.", "score": 50.0}404{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_95.wav", "doc_id": "uZBWfYjYnf.seg_95", "src_text": "And what are the problems of the current SimulST models?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "sind die Probleme der aktuellen Simulink-Modelle?", "score": 50.0}405{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_215.wav", "doc_id": "oYCKgTzTDy.seg_215", "src_text": "Hello everyone, my name is Yusen Zhang from the Penn State University.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hallo alle, mein Name ist Usain John von der Universität von Pennsylvania.", "score": 71.0}406{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_67.wav", "doc_id": "TVCREhgqUP.seg_67", "src_text": "For the first time, we show strong generalization to deeper recursion without relying on trees.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum ersten Mal zeigen wir eine starke Generalisierung zu tiefer Rekursion ohne auf Bäume zurückgreifen zu müssen.", "score": 70.0}407{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_789.wav", "doc_id": "WTTtiRKFZI.seg_789", "src_text": "So \"Marge read it yesterday\" is fine because the direct object is close to the verb, while \"Marge read yesterday it\" is much worse.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "is close to the verb.“ Während es gestern Abend viel schlimmer war, ist es heute viel besser, weil", "score": 5.0}408{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_856.wav", "doc_id": "GvEBWkLmuI.seg_856", "src_text": "However, when we actually look at the distribution of the words and lexicon, we find very different things.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wir verwenden. Wir finden jedoch sehr unterschiedliche Dinge, wenn wir uns die Verteilung der Wörter im Lexikon ansehen. Also haben", "score": 100.0}409{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_421.wav", "doc_id": "WBLMIsdIrq.seg_421", "src_text": "But then if we use COMET, context-aware models perform best.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "haben. Dann verwenden wir jedoch Comet, kontextbewusste Modelle, die am besten funktionieren,", "score": 100.0}410{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_519.wav", "doc_id": "dvGkKzmIaN.seg_519", "src_text": "Then let me introduce the details of our embedding marker.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dann lassen Sie mich die Einzelheiten unseres eingebetteten Markers erläutern.", "score": 85.0}411{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_362.wav", "doc_id": "gGbuDbHhyc.seg_362", "src_text": "This indicates that WSL approaches actually require cleanly labeled data to work properly, and the annotation cost for obtaining clean validation samples should not be overlooked.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies zeigt, dass WSL-Ansätze tatsächlich sauber gekennzeichnete Daten erfordern, um ordnungsgemäß zu funktionieren, und dass die Anmerkungskosten für die Erlangung sauberer Validierungsbeispiele nicht vernachlässigt werden sollten.", "score": 95.0}412{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_361.wav", "doc_id": "gGbuDbHhyc.seg_361", "src_text": "As shown in this figure, if there are no clean validation samples, then the trained models cannot generalize beyond the original weak labels, meaning that the training is pointless.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wie in dieser Abbildung zu sehen ist. Wenn es keine sauberen Validierungsmuster gibt, können die Trendmodelle nicht über die ursprünglichen Bit-Etiketten hinaus generalisiert werden. Das bedeutet, dass diese Doktrin sinnlos ist.", "score": 50.0}413{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_631.wav", "doc_id": "oeooqChmKK.seg_631", "src_text": "Thanks for listening.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "für das Zuhören.", "score": 65.0}414{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_751.wav", "doc_id": "XejEJmgUmE.seg_751", "src_text": "Finally, we can choose sentences from a completely unrelated domain such as Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Schließlich können wir Sätze aus einer völlig unabhängigen Domäne wie Wikipedia auswählen.", "score": 100.0}415{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_508.wav", "doc_id": "dvGkKzmIaN.seg_508", "src_text": "However, recent works have shown that the attacker may steal the model through learning from the embedding and provide similar services.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Arbeiten gezeigt, dass der Angreifer das Modell durch das Lernen vom Embedding stehlen und ähnliche Dienste anbieten kann.", "score": 83.0}416{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_783.wav", "doc_id": "WTTtiRKFZI.seg_783", "src_text": "So we get dependencies from the governor.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "so dass wir von der Regierungsstruktur Abhängigkeiten erhalten, die", "score": 75.0}417{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_853.wav", "doc_id": "GvEBWkLmuI.seg_853", "src_text": "So for instance, for the personas of black women, we would do Fightin’ Words and compare the log-odds ratios against both white personas and man personas because those are the two corresponding unmarked groups.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zum Beispiel für die Persönlichkeiten der schwarzen Frauen werden wir Wörter sammeln und die Logod-Ratios für beide weißen und männlichen Persönlichkeiten vergleichen, weil es sich um zwei nicht markierte Gruppen handelt.", "score": 53.0}418{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_277.wav", "doc_id": "PIZEXUFLAR.seg_277", "src_text": "We follow the method from OFA and formulate all the tasks in a unified sequence-to-sequence format.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "folgen wir der Masterform von OFA und formulieren alle Aufgaben in einer vereinheitlichten Sequenz-zu-Sequenz-Format.", "score": 74.0}419{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_701.wav", "doc_id": "oaOHnMCwad.seg_701", "src_text": "We host 2 tasks on lab in the wild, one of them being social acceptability, and the way this works is that participants will read a situation from the social chemistry dataset and, then they'll write how socially acceptable a situation is.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und die Art und Weise, wie diese Arbeit ist, dass die Teilnehmer eine Situation aus der Sozialchemie erhalten und wie sozialverträglich diese Situation ist.", "score": 50.0}420{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_546.wav", "doc_id": "rISrKoXQCx.seg_546", "src_text": "Hi, I'm Shangbin, PhD student in the University of Washington.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hallo, ich bin John Bin, Doktorand an der University of Washington.", "score": 60.0}421{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_255.wav", "doc_id": "oYCKgTzTDy.seg_255", "src_text": "For example, Encoder-Decoder outperforms previous work or achieves comparable results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "zum Beispiel Encoder-Decoder-Modelle, die vergleichbare Ergebnisse", "score": 71.0}422{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_282.wav", "doc_id": "PIZEXUFLAR.seg_282", "src_text": "We use all the instances in the test split for each task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Prüfung aus. Wir verwenden alle Instanzen im Test für", "score": 82.0}423{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_747.wav", "doc_id": "XejEJmgUmE.seg_747", "src_text": "And we can also do the same by choosing sentences from a different subset or a different data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und wir können dasselbe auch erreichen, indem wir Sätze aus einem anderen Satz oder einem anderen Datensatz auswählen.", "score": 93.0}424{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_52.wav", "doc_id": "TVCREhgqUP.seg_52", "src_text": "As usual, we have a training set of utterances.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "aus, als hätten Sie in diesem Fall ein", "score": 0.0}425{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_100.wav", "doc_id": "uZBWfYjYnf.seg_100", "src_text": "So what is our solution?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Was ist unsere Lösung?", "score": 96.0}426{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_844.wav", "doc_id": "GvEBWkLmuI.seg_844", "src_text": "Our prompts to generate these personas were inspired by a study where they gave these prompts to human subjects, finding that by giving it to human subjects, they also were able to surface racial stereotypes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unsere Anfragen zur Erzeugung dieser Personen wurden von einer Studie inspiriert, in der sie diesen Anfragen menschliche Teilnehmer gaben und fanden, dass sie auch Rassenstereotypen aufdecken konnten.", "score": 91.0}427{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_246.wav", "doc_id": "oYCKgTzTDy.seg_246", "src_text": "We found that Encoder-Decoder or Encoder-PTR can be improved by training in a mixture of various languages.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir stellten fest, dass Encoder-Decoder oder Encoder-PDR durch Schulung in einer Mischung verschiedener Sprachen verbessert werden können.", "score": 100.0}428{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_667.wav", "doc_id": "FLkGnzVRew.seg_667", "src_text": "We find that PRC has the highest percentage of dissonance and works best for rare class.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir stellen fest, dass die Anmerkung mit dem höchsten Prozentsatz von Anmerkungen und der besten Anmerkung für die Klasse ist.", "score": 27.0}429{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_515.wav", "doc_id": "dvGkKzmIaN.seg_515", "src_text": "Finally, the watermark needs to be transferable to the attacker's services during the model extraction process.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "schließlich muss das Wasserzeichen während des Modellextraktionsprozesses auf die Services des Angreifers übertragbar sein.", "score": 96.0}430{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_182.wav", "doc_id": "SLpqvupgvW.seg_182", "src_text": "We provide the first and second speech bubbles automatically, but the third one is filled in by the annotator.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir stellen automatisch die ersten beiden Sprachbubbles zur Verfügung, aber der dritte wird vom Kommentator gefüllt.", "score": 96.0}431{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_859.wav", "doc_id": "GvEBWkLmuI.seg_859", "src_text": "And in fact, this lexicon doesn't really capture many of the harmful patterns that we saw in the earlier slides well at all.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und in der Tat hat der Lesekontext nicht wirklich viele der schädlichen Muster erfasst, die wir in den früheren Zeilen gesehen haben,", "score": 82.0}432{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_877.wav", "doc_id": "GvEBWkLmuI.seg_877", "src_text": "We just really can't make any assumptions or really study that further, without more transparency.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir konnten wirklich keine Annahmen treffen, und wir untersuchten das weiter mit mehr Transparenz.", "score": 65.0}433{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_582.wav", "doc_id": "rISrKoXQCx.seg_582", "src_text": "So if we do not sanitize political opinions in language model training data, the bias would propagate from pretraining data to language models to downstream tasks, ultimately creating fairness issues.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wenn wir die politischen Meinungen in den Sprachtrainingsdaten nicht standardisieren, würde sich die Vorliebe auf die vorherigen Sprachmodelle bis hinunter zu den Aufgaben erstrecken, was letztlich zu Fairnessfragen führen würde.", "score": 75.0}434{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_462.wav", "doc_id": "hgIDlKNiFM.seg_462", "src_text": "So thank you for this presentation, and we are looking forward to exchange at the poster session in Toronto.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dank für diese Präsentation, und wir freuen uns auf Aktionen bei der Post in Toronto.", "score": 57.0}435{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_192.wav", "doc_id": "SLpqvupgvW.seg_192", "src_text": "The third one is when they have similar descriptions on Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und die dritte Methode ist die gleichartige Zufallsauswahl, z.B. zwei Bücher", "score": 64.0}436{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_171.wav", "doc_id": "SLpqvupgvW.seg_171", "src_text": "Here are some examples of indirect references for example, \"the newer one\" or \"the song that's not energetic.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "hier sind einige Beispiele für direkte Präferenzen, z. B. der neueste oder der nicht energiereiche.", "score": 69.0}437{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_134.wav", "doc_id": "wLqFAuDnKa.seg_134", "src_text": "The majority of sentences 516 out of 1,000.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die Mehrheit der Sätze – 516 von 1000", "score": 100.0}438{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_4.wav", "doc_id": "aQpIWggfCo.seg_4", "src_text": "And show that large language models can effectively decompose goals into steps.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und zeigen, dass große Sprachmodelle Ziele effektiv in Schritte zerlegen können.", "score": 98.0}439{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_677.wav", "doc_id": "oaOHnMCwad.seg_677", "src_text": "So let's start off by imagining that you're working for a newspaper and you're sifting through comments under your news article trying to remove toxic content.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Lassen Sie uns also davon ausgehen, dass Sie für eine Zeitung arbeiten und versuchen Sie, den Inhalt zu entfernen.", "score": 57.0}440{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_629.wav", "doc_id": "oeooqChmKK.seg_629", "src_text": "Still, even the best-performing models seem to have difficulties with reliably integrating backward knowledge presented only at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Stille. Auch die besten Modelle scheinen Schwierigkeiten mit zuverlässig integriertem rückwärts gerichteten Wissen zu haben, das nur bei der Inferenzzeit vorgestellt wird.", "score": 91.0}441{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_424.wav", "doc_id": "WBLMIsdIrq.seg_424", "src_text": "Now, we use the MuDA benchmark to evaluate models and we find that context-aware models are significantly more accurate than models that do not use context for certain discourse phenomena such as formality and lexical cohesion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "verwenden wir den Munda-Benchmark, um Modelle zu bewerten, und wir stellen fest, dass Kontextmodellen signifikant genau sind als Modelle, die Kontext nicht für", "score": 88.0}442{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_578.wav", "doc_id": "rISrKoXQCx.seg_578", "src_text": "So this has sound the alarm for us to acknowledge and tackle the fairness issues resulting by language model political leanings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "sich greifen könnte. Diese klingen also wie eine Warnung für Sie, um die Fairness-Probleme zu erkennen und anzugehen,", "score": 55.0}443{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_54.wav", "doc_id": "TVCREhgqUP.seg_54", "src_text": "And \"Mary knew that the girl slept.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "ein neues Trainingsprogramm für die Mädchen. Dies", "score": 2.0}444{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_351.wav", "doc_id": "gGbuDbHhyc.seg_351", "src_text": "Technically, this claim is not wrong, but there's a catch, which is that people do assume that there's an additional clean validation set available for model selection.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Tatsächlich ist diese Behauptung nicht falsch, aber es gibt einen Haken. Denn die Leute gehen davon aus, dass es für die Modellauswahl ein zusätzliches sauberes", "score": 98.0}445{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_30.wav", "doc_id": "aQpIWggfCo.seg_30", "src_text": "Since large language models are costly to deploy, it's essential to enable language planning ability of smaller and specialized models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Da große Sprachmodelle teuer zu implementieren sind, ist es unerlässlich, Sprachplanung für kleinere und spezialisierte Modelle zu ermöglichen.", "score": 97.0}446{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_220.wav", "doc_id": "oYCKgTzTDy.seg_220", "src_text": "Existing cross-lingual semantic parsing models are separately proposed and evaluated on data set of limited tasks and applications.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Bestehende Cross-Lingual-Semantic-Parsing-Modelle werden separat auf Datensätzen mit begrenzten Aufgaben und Anwendungen vorgeschlagen", "score": 90.0}447{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_254.wav", "doc_id": "oYCKgTzTDy.seg_254", "src_text": "We also find some other interesting findings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir werden auch einige andere interessante Erkenntnisse", "score": 92.0}448{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_677.wav", "doc_id": "oaOHnMCwad.seg_677", "src_text": "So let's start off by imagining that you're working for a newspaper and you're sifting through comments under your news article trying to remove toxic content.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "durchgeführt. Jetzt ist klar, dass Sie für eine Zeitung arbeiten und Kommentare und Artikel schreiben und versuchen, den Inhalt zu entfernen.", "score": 26.0}449{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_319.wav", "doc_id": "dJGfOSFgZO.seg_319", "src_text": "We call this approach annotating behaviors in chat or ABC-Eval in short.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir nennen diesen Ansatz „Annotieren von Verhaltensweisen in Chats“ oder „ABC-Eval in Kürze“.", "score": 68.0}450{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_101.wav", "doc_id": "uZBWfYjYnf.seg_101", "src_text": "First, to use already existing offline ST models without re-training or adopting specific architecture for SimulST.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Erstens, verwenden Sie bereits vorhandene Online-SD-Modelle ohne Wiedertrainieren oder die Anpassung spezifischer Architekturen für CivilSD; verwenden", "score": 58.0}451{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_416.wav", "doc_id": "WBLMIsdIrq.seg_416", "src_text": "And we called our tagger the Multilingual Discourse-Aware, or MuDA tagger.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "nennen unseren Tigrer den multilingualen Diskurs bewusst oder umuda Tigrer.", "score": 48.0}452{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_341.wav", "doc_id": "gGbuDbHhyc.seg_341", "src_text": "Hello, I am Dawei, a PhD student at Saarland University in Germany.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo, ich bin Dawid, ein Doktorand an der Universität Salent in Deutschland.", "score": 65.0}453{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_241.wav", "doc_id": "oYCKgTzTDy.seg_241", "src_text": "And we also find many interesting results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir stellen auch viele interessante Ergebnisse", "score": 91.0}454{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_29.wav", "doc_id": "aQpIWggfCo.seg_29", "src_text": "Our method greatly improves the planning ability both in semantic completeness and faithfulness to the constraint.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Unsere Methode verbessert die Anfälligkeit sowohl in semantischer Vollständigkeit als auch in Treue zu den Einschränkungen.", "score": 87.0}455{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_489.wav", "doc_id": "SUkmfOTvGi.seg_489", "src_text": "And this shows us that adaptive overfitting in this case is not observed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und das zeigt uns, dass adaptive Überanpassung in diesem Fall nicht beobachtet wird.", "score": 100.0}456{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_509.wav", "doc_id": "dvGkKzmIaN.seg_509", "src_text": "Therefore, it's necessary to protect the copyright of embedding as services.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Daher ist es notwendig, das Urheberrecht von Embedding- und Services zu schützen.", "score": 74.0}457{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_327.wav", "doc_id": "dJGfOSFgZO.seg_327", "src_text": "In addition, ABC-Eval labels are more predictive of the overall conversation quality compared to metrics produced by existing methods, as shown by this simple linear regression analysis.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "zu Metriken, die durch existierende Methoden erzeugt werden, wie durch die einfache Regressionsanalyse. Beispielsweise können Sie sehen, wie die Proportionen der Drehungen mit", "score": 94.0}458{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_783.wav", "doc_id": "WTTtiRKFZI.seg_783", "src_text": "So we get dependencies from the governor.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Gouverneurs hier die Abhängigkeiten von allen", "score": 81.0}459{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_373.wav", "doc_id": "gGbuDbHhyc.seg_373", "src_text": "Their performance gain and practicality are heavily overestimated.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wobei Leistungszuwächse und Praktikabilität stark überschätzt werden.", "score": 100.0}460{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_467.wav", "doc_id": "SUkmfOTvGi.seg_467", "src_text": "We observe that models have been used in CoNLL-2003 to develop NER for almost 20 years and this naturally raises several problems.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir beobachten, dass Modelle seit 2003 Kornell verwendet haben, um NER zu entwickeln. Das wirft natürlich einige Probleme auf.", "score": 60.0}461{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_712.wav", "doc_id": "oaOHnMCwad.seg_712", "src_text": "We also find most additional alignment with people who have a college education.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir finden auch die meisten zusätzlichen Zuordnungen zu Personen mit Hochschulbildung, daher finden", "score": 90.0}462{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_35.wav", "doc_id": "aQpIWggfCo.seg_35", "src_text": "In total, we generate 55,000 specific goals with scripts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Insgesamt generieren wir fünfzigtausend spezifische Ziele mit Skripten,", "score": 65.0}463{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_396.wav", "doc_id": "WBLMIsdIrq.seg_396", "src_text": "And second, how well do models handle these cases?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und zweitens, wie gut können die Modelle diese Fälle handhaben.", "score": 100.0}464{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_420.wav", "doc_id": "WBLMIsdIrq.seg_420", "src_text": "First of all, when we use corpus-level metrics: so for BLEU, we find that context-agnostic models have the best performance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Erstens, wenn wir Korpus-Level-Metriken verwenden, sehen wir, dass die komplexen agnostischen Modelle die beste Leistung", "score": 60.0}465{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_372.wav", "doc_id": "gGbuDbHhyc.seg_372", "src_text": "To summarize, we showed that recent WSL approaches require clean, manually annotated samples for them to work properly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zusammenfassend lässt sich sagen, dass aktuelle WSL-Ansätze saubere, manuell annotierte Proben benötigen, um richtig zu funktionieren,", "score": 100.0}466{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_358.wav", "doc_id": "gGbuDbHhyc.seg_358", "src_text": "We addressed these research questions in our work and our findings are as follows.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir gehen in unserer Arbeit auf diese Forschungsfragen ein, und unsere Ergebnisse sind wie folgt.", "score": 100.0}467{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_756.wav", "doc_id": "XejEJmgUmE.seg_756", "src_text": "And we saw here in the orange dotted line, the MPP judgments are relatively stable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir sahen hier in der Orange-Dot-Zeile, dass die MP-P-Juristen relativ stabil sind.", "score": 90.0}468{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_26.wav", "doc_id": "aQpIWggfCo.seg_26", "src_text": "In addition, we reward the script that contains the keywords of the target constraint.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Darüber hinaus vermeiden wir das Skript, das die Schlüsselwörter der Zielbeschränkung enthält.", "score": 91.0}469{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_657.wav", "doc_id": "FLkGnzVRew.seg_657", "src_text": "Thus, this is the model that we use to cold start the active learning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "mit dem wir beginnen, viel besser ist.", "score": 60.0}470{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_438.wav", "doc_id": "hgIDlKNiFM.seg_438", "src_text": "Since its release in 2018, BERT has become one of the most effective approach to solve natural language processing tasks and offers huge performance gains compared to historical static and contextualized methods such as Word2vec, fastText, or more.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Seit seiner Veröffentlichung im Dezember ist BERT zu einem der effektivsten Ansätze für die Verarbeitung natürlicher Sprache geworden und bietet im Vergleich zu historischen statischen und kontextualisierten Methoden enorme Leistungssteigerungen.", "score": 93.0}471{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_582.wav", "doc_id": "rISrKoXQCx.seg_582", "src_text": "So if we do not sanitize political opinions in language model training data, the bias would propagate from pretraining data to language models to downstream tasks, ultimately creating fairness issues.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir politische Meinungen in Sprachmodellierungstrainingdaten nicht saniert haben, würde der Bias sich von vorbereitenden Daten zu Sprachmodellen und schließlich zu Downstream-Aufgaben ausbreiten, was letztendlich zu Fairnessproblemen führen würde.", "score": 81.0}472{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_55.wav", "doc_id": "TVCREhgqUP.seg_55", "src_text": "These utterances are paired with logical forms that represent core aspects of their meaning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "hatte. Diese Aussagen werden mit logischen Formen gepaart, die die Kernaspekte ihres Sinns darstellen.", "score": 95.0}473{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_612.wav", "doc_id": "oeooqChmKK.seg_612", "src_text": "First, we have the typical setting: \"Background-Pretrain\", where background knowledge is assumed to be available at pretrain time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Erstens den typischen Einstellung: Hintergrund vorbereiten, bei der man davon ausgeht, dass Hintergrundwissen zur Vorbereitung verfügbar ist.", "score": 65.0}474{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_762.wav", "doc_id": "XejEJmgUmE.seg_762", "src_text": "So why does the match prefix affect the language model judgement so much?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "beeinflussen. Daher, warum beeinflusst der Match-Präfix die Sprachmodellbewertung so stark?", "score": 100.0}475{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_258.wav", "doc_id": "oYCKgTzTDy.seg_258", "src_text": "We conduct a comprehensive benchmark study on three representative types of multilingual language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir führen eine umfassende Benchmark-Studie zu drei repräsentativen Typen von mehrsprachigen Sprachmodellen durch,", "score": 98.0}476{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_712.wav", "doc_id": "oaOHnMCwad.seg_712", "src_text": "We also find most additional alignment with people who have a college education.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir finden auch die meisten zusätzlichen Angaben zu Personen mit Hochschulbildung, so dass", "score": 84.0}477{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_120.wav", "doc_id": "uZBWfYjYnf.seg_120", "src_text": "And we also released open source the code and models and simultaneous output to facilitate the reproducibility of our work.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "und wir stellen auch Open-Source-Code und Modelle zur Verfügung, um die Reproduzierbarkeit unserer Arbeit zu erleichtern,", "score": 90.0}478{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_518.wav", "doc_id": "dvGkKzmIaN.seg_518", "src_text": "Therefore, in this paper we propose Embedding marker, which is a backdoor based watermark method applicable to embedding as services.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daher schlagen wir in diesem Papier Embedder vor, eine Backdoor-basierte Methode zur Anwendung auf Embedding- und Dienstleistungen.", "score": 75.0}479{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_581.wav", "doc_id": "rISrKoXQCx.seg_581", "src_text": "It's like between Scylla and Charybdis.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "werden, wie zwischen s und kris.", "score": 22.0}480{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_693.wav", "doc_id": "oaOHnMCwad.seg_693", "src_text": "Our framework works in two main steps.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Unser Rahmenwerk funktioniert in zwei Hauptschritten.", "score": 100.0}481{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_812.wav", "doc_id": "WTTtiRKFZI.seg_812", "src_text": "So the proportion is bigger of the left short conjunct.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist die Proportion größer von dem linken kürzeren Konjunkt,", "score": 95.0}482{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_736.wav", "doc_id": "XejEJmgUmE.seg_736", "src_text": "And then the hope is that the model, basically, puts more probability to the acceptable sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "die Hoffnung, dass das Modell im Grunde mehr Wahrscheinlichkeit auf die akzeptable Situation legt.", "score": 85.0}483{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_114.wav", "doc_id": "uZBWfYjYnf.seg_114", "src_text": "And we compare with popular strategies that are also applied to offline models that are the Wait-k strategy and the Local Agreement.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und wir vergleichen mit geeigneten Strategien, die auch auf Offline-Modelle angewendet werden können, nämlich die Whitkey-Strategie und die lokale Vereinbarung,", "score": 60.0}484{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_282.wav", "doc_id": "PIZEXUFLAR.seg_282", "src_text": "We use all the instances in the test split for each task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Milleisenas auswählen. Wir verwenden alle Instanzen im Testset für", "score": 63.0}485{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_403.wav", "doc_id": "WBLMIsdIrq.seg_403", "src_text": "And we perform our analysis on transcripts of TED talks that have been translated from English to 14 different languages.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und wir führen unsere Analysen auf Transkripte von Ted Talks durch, die in vierzehn verschiedenen Sprachen übersetzt wurden.", "score": 99.0}486{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_868.wav", "doc_id": "GvEBWkLmuI.seg_868", "src_text": "And finally, for black women, we see that some of the top words are things like \"strong\" and \"resilient\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "schließlich sehen wir bei den schwarzen Frauen, dass einige der Top-Wörter Dinge wie stark und widerstandsfähig sind.", "score": 92.0}487{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_505.wav", "doc_id": "dvGkKzmIaN.seg_505", "src_text": "Currently, large language models such as GPT, LLAMA, PALM are exceptional in natural language understanding and generation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "zunächst die Hintergründe der Einbettung von Diensten erläutern. Derzeit sind große Sprachmodelle wie TpT, Llama, Palm im Bereich des natürlichen Sprachverständnisses und der Sprachgenerierung", "score": 13.0}488{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_285.wav", "doc_id": "PIZEXUFLAR.seg_285", "src_text": "During training, we mix all the instances for all the tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "während des Trainings mischen wir alle Instanzen für alle Aufgaben,", "score": 98.0}489{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_401.wav", "doc_id": "WBLMIsdIrq.seg_401", "src_text": "We can think of words that have high P-CXMI as ones that require context for translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir können annehmen, dass Wörter mit hohem XMI eine Übersetzung erfordern.", "score": 97.0}490{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_621.wav", "doc_id": "oeooqChmKK.seg_621", "src_text": "We evaluate the data set both with human study participants, and established coreference resolution models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir bewerten das Datensatz sowohl mit menschlichen Studienteilnehmern als auch mit etablierten Frage-Antwort-Modellen.", "score": 59.0}491{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_363.wav", "doc_id": "gGbuDbHhyc.seg_363", "src_text": "Our second finding is that increasing the number of clean validation samples will help WSL approaches to achieve better performance, as shown in the figure on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Unsere zweite Erkenntnis ist, dass die Erhöhung der Anzahl der Reinheitsvalidierungsproben den WSS-Ansätzen helfen wird, eine bessere Leistung zu erzielen, wie in der Abbildung auf der linken Seite dargestellt.", "score": 93.0}492{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_285.wav", "doc_id": "PIZEXUFLAR.seg_285", "src_text": "During training, we mix all the instances for all the tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "während des Trainings vermischen wir alle Instanzen für alle Aufgaben;", "score": 100.0}493{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_36.wav", "doc_id": "aQpIWggfCo.seg_36", "src_text": "To ensure the quality of the validation and test set, we ask crowd-sourced workers to find and revise the incorrect samples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "um die Qualität der Validierung und Teststätten zu gewährleisten, und bitten Crowdsource-Worker, die unkorrekten Muster zu finden und zu korrigieren.", "score": 97.0}494{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_757.wav", "doc_id": "XejEJmgUmE.seg_757", "src_text": "Now, what happens when we choose sentences from the same data set?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Was passiert nun, wenn wir Sätze aus demselben Datensatz auswählen?", "score": 97.0}495{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_295.wav", "doc_id": "PIZEXUFLAR.seg_295", "src_text": "Also, transfer learning from natural instruction dataset can benefit instruction tuning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "natürlichen Sprachdatensätzen kann beim Anpassen von Anweisungen helfen.", "score": 65.0}496{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_530.wav", "doc_id": "dvGkKzmIaN.seg_530", "src_text": "Copyright verification is to detect whether a model behind another service contains the word mark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die Urheberrechtsprüfung soll feststellen, ob ein Modell hinter einem anderen Dienst die Wasserzeichen enthält.", "score": 62.0}497{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_567.wav", "doc_id": "rISrKoXQCx.seg_567", "src_text": "We separately pretrain language models on the two different temporal corpora.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und trennen Sprachmodelle auf zwei verschiedene zeitliche Korpora.", "score": 92.0}498{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_740.wav", "doc_id": "XejEJmgUmE.seg_740", "src_text": "We're trying to revisit the MPP pipeline by asking the model to evaluate acceptability on longer and longer sequences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "versuchen, indem wir bitten, das Modell zu überprüfen, um die Akzeptanz auf längere und längere Sequenzen zu bewerten.", "score": 64.0}499{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_174.wav", "doc_id": "SLpqvupgvW.seg_174", "src_text": "Our data set covers three different domains: music, books, and recipes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "unser Datensatz umfasst drei verschiedene Bereiche: Musik, Bücher und Rezepte.", "score": 100.0}500{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_384.wav", "doc_id": "WBLMIsdIrq.seg_384", "src_text": "A Data-driven, Multilingual Exploration\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "vorstellen.", "score": 0.0}501{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_817.wav", "doc_id": "WTTtiRKFZI.seg_817", "src_text": "Here we have coordination of two verbs and there's no outsides, external governor.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Außenstelle des Gouverneurs nicht zusteht, die beiden Länder gegeneinander auszuspielen.", "score": 0.0}502{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_91.wav", "doc_id": "TVCREhgqUP.seg_91", "src_text": "If you want to learn more about our experiments and how we address these challenges, please have a look at our paper or come to our poster.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn Sie mehr über unsere Experimente und wie wir diese Herausforderungen angehen möchten, bitten wir Sie, unsere Publikation oder unsere Poster zu besuchen.", "score": 83.0}503{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_28.wav", "doc_id": "aQpIWggfCo.seg_28", "src_text": "With our method, InstructGPT can generate scripts of higher quality.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Mit unserer Methode kann Insensitivität zu höherer Qualität führen.", "score": 30.0}504{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_684.wav", "doc_id": "oaOHnMCwad.seg_684", "src_text": "Positionality is simply the perspectives that people hold as a result of their demographics, identity, and life experiences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Positionierung ist einfach die Perspektiven, die Menschen aufgrund ihrer Demografie, Identität und Lebenserfahrungen haben.", "score": 82.0}505{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_458.wav", "doc_id": "hgIDlKNiFM.seg_458", "src_text": "Which is not the case for the model based on CamemBERT weights and tokenizer, which suffer from stability issues.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies ist nicht der Fall für das Modell, das auf Camberweights und Token basiert, die aus Stabilitätsgründen stammen.", "score": 30.0}506{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_368.wav", "doc_id": "gGbuDbHhyc.seg_368", "src_text": "Finally, the performance improvement claimed in previous WSL approaches can be easily achieved by allowing to continue fine-tuning on the clean validation samples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Schließlich kann die Leistung, die in früheren WS-L-Annäherungen behauptet wurde, leicht erreicht werden, indem man weiterhin feine Anpassungen auf sauberen Validierungssampeln vornimmt.", "score": 50.0}507{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_646.wav", "doc_id": "FLkGnzVRew.seg_646", "src_text": "We used dissonance-first approach, as seen in the flow chart here.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir verwenden den Distanz-First-Ansatz, wie in der Flowchart hier zu sehen ist. „Tweets“ werden mit einem", "score": 60.0}508{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_237.wav", "doc_id": "oYCKgTzTDy.seg_237", "src_text": "And during inference we can use this model to translate German queries or Chinese queries, et cetera.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "und können dieses Modell während des Lernens verwenden. Um deutsche oder chinesische Anfragen zu übersetzen. Und wir berücksichtigen", "score": 55.0}509{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_709.wav", "doc_id": "oaOHnMCwad.seg_709", "src_text": "For example, we find that data sets and models are most aligned to English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Beispiel finden wir heraus, dass die Datenmodelle für die meisten englischsprachigen Länder am besten geeignet sind,", "score": 99.0}510{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_647.wav", "doc_id": "FLkGnzVRew.seg_647", "src_text": "Tweets were passed using the PDTB parser, and pairs of discourse units were annotated according to the guidelines that are described in our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "„PDT-Parser“ verarbeitet und „Paare von Diskussionsgruppen“ werden entsprechend der in der Leitlinie beschriebenen Anweisungen annotiert.", "score": 85.0}511{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_848.wav", "doc_id": "GvEBWkLmuI.seg_848", "src_text": "So the Marked Words method draws upon the sociolinguistic concept of \"markedness\", which states that there is an unmarked default, and any group that differs from that default is linguistically marked.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Markierungsmethode bezieht sich auf den soziolinguistischen Begriff der Markiertheit, der besagt, dass es sich um eine unmarkierte Gruppe handelt. Also zum Beispiel das Wort Mann oder Krieger ist normalerweise mit dem", "score": 50.0}512{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_439.wav", "doc_id": "hgIDlKNiFM.seg_439", "src_text": "Since then, this model has been adapted to many other languages, like in French with CamemBERT, and also in domains like biomedical with PubMedBERT and BioBERT and on clinical with ClinicalBERT, but mostly in English.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Seitdem wurde dieses Modell auf viele andere Sprachen wie Französisch mit CamemBERT, andere Domänen wie Biomedizin mit PubMedBERT und BioBERT, und klinische mit ClinicalBERT, aber hauptsächlich auf Englisch adaptiert.", "score": 99.0}513{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_540.wav", "doc_id": "dvGkKzmIaN.seg_540", "src_text": "We also validate the covertness of the provided embedding by visualising the embedding of sentences on four dataset [INAUDIBLE 4:39] PCA.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir haben auch die Geheimniskrämerei der vorgestellten Einbettung durch die Visualisierung der Einbettung von Sätzen auf 40. z. v. p. c. A. bestätigt,", "score": 50.0}514{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_740.wav", "doc_id": "XejEJmgUmE.seg_740", "src_text": "We're trying to revisit the MPP pipeline by asking the model to evaluate acceptability on longer and longer sequences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir versuchen, die MPB-Pipeline zu überprüfen, indem wir die Modelle bitten, die Akzeptanz auf längere und längere Sequenzen zu bewerten.", "score": 65.0}515{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_419.wav", "doc_id": "WBLMIsdIrq.seg_419", "src_text": "And finally, we use our benchmark as well as other metrics to evaluate different models on the document-level machine translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und schließlich verwenden wir unseren Benchmark wie andere Matrizen, um verschiedene Modelle auf der Dokumenten-Ebene der Maschinentranslation zu bewerten.", "score": 70.0}516{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_852.wav", "doc_id": "GvEBWkLmuI.seg_852", "src_text": "So in our method, we first designate what the unmarked and marked groups are, and then we compare the personas using the Fightin’ Words method, which is basically using weighted log-odds ratios to distinguish the top words for each marked group.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In unserer Methode werden zuerst die unmarkierten und markierten Gruppen ermittelt. Und dann vergleichen wir die Personen, die die Fighting Words-Methode verwenden, die im Grunde die Weighted Logit-Methode verwendet, um die Top-Wörter jeder Gruppe zu unterscheiden.", "score": 80.0}517{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_673.wav", "doc_id": "FLkGnzVRew.seg_673", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Danke.", "score": 100.0}518{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_462.wav", "doc_id": "hgIDlKNiFM.seg_462", "src_text": "So thank you for this presentation, and we are looking forward to exchange at the poster session in Toronto.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "danken wir Ihnen für diese Präsentation und wir freuen uns darauf, in Toronto bei der Postzustellung Aktionen zu unternehmen.", "score": 58.0}519{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_717.wav", "doc_id": "oaOHnMCwad.seg_717", "src_text": "So, given that there is positionality in NLP, what can we do about it?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Angesichts der Tatsache, dass es sich um eine Position in einer LED und LP handelt, was können wir tun?", "score": 59.0}520{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_395.wav", "doc_id": "WBLMIsdIrq.seg_395", "src_text": "First, when does translation require context?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Erstens: Wann benötigt eine Übersetzung einen", "score": 100.0}521{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_633.wav", "doc_id": "FLkGnzVRew.seg_633", "src_text": "I would like to present our work accepted into ACL 2023 as a long paper, \"Transfer Learning for Dissonance Detection: Addressing the Rare-Class Challenge.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Ich würde gerne unsere Arbeit, die in den ACL 23 als langen Papier zur Transfer-Learning für die Erkennung von Dissonanzdetektionen, die sich der seltenen Klasse widmet, angenommen wurde, präsentieren.", "score": 40.0}522{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_593.wav", "doc_id": "oeooqChmKK.seg_593", "src_text": "But natural language understanding often requires knowledge that is also supplied at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Aber das Verständnis der natürlichen Sprache erfordert oft Wissen, das auch im Zeitraum der Nachsorge bereitgestellt wird.", "score": 70.0}523{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_29.wav", "doc_id": "aQpIWggfCo.seg_29", "src_text": "Our method greatly improves the planning ability both in semantic completeness and faithfulness to the constraint.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Unser Verfahren verbessert die Schmerztoleranz sowohl in der semantischen Vollständigkeit als auch in der Treue gegenüber den Einschränkungen.", "score": 88.0}524{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_386.wav", "doc_id": "WBLMIsdIrq.seg_386", "src_text": "So a lot of translations depend on context.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Viele Übersetzungen hängen vom Kontext ab, zum Beispiel:", "score": 79.0}525{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_163.wav", "doc_id": "SLpqvupgvW.seg_163", "src_text": "Consider this alternative question.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "möchten. Überlegen Sie sich diese alternative", "score": 90.0}526{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_737.wav", "doc_id": "XejEJmgUmE.seg_737", "src_text": "The current MPP pipeline basically doesn't allow us to evaluate a model's acceptance towards longer sentences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die aktuelle MPP-Pipeline ermöglicht es uns im Grunde nicht, die Akzeptanz eines Modells für längere Sätze zu bewerten.", "score": 100.0}527{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_73.wav", "doc_id": "TVCREhgqUP.seg_73", "src_text": "This makes our approach quite flexible and expressive.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "hat. Dies macht unseren Ansatz sehr flexibel und ausdrucksstark.", "score": 100.0}528{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_454.wav", "doc_id": "hgIDlKNiFM.seg_454", "src_text": "However, we can observe that data from heterogeneous sources appear to be more versatile.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir können jedoch feststellen, dass Daten aus heterogenen Quellen vielseitiger zu sein scheinen,", "score": 100.0}529{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_245.wav", "doc_id": "oYCKgTzTDy.seg_245", "src_text": "And we evaluate on mT5 and XLM-R + PTR on multilingual setting.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und wir evaluieren das auf Mf fünf und das Beispiel Xlm plus Pd auf mehrsprachige Einstellungen.", "score": 30.0}530{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_202.wav", "doc_id": "SLpqvupgvW.seg_202", "src_text": "For example, the one with the piano music.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zum Beispiel der mit der Klaviermusik.", "score": 95.0}531{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_668.wav", "doc_id": "FLkGnzVRew.seg_668", "src_text": "However, the annotators also find the examples difficult.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Annotatoren stellen auch heraus, dass die Beispiele schwierig sind.", "score": 85.0}532{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_105.wav", "doc_id": "uZBWfYjYnf.seg_105", "src_text": "Our solution is to propose EDAtt, or Encoder-Decoder Attention, and it is a strategy for which we decide whether to emit or not a partial translation, based on where attention points to.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Unsere Lösung besteht darin, die Aufmerksamkeit zu fokussieren oder zu kodieren, und es ist eine Strategie, bei der wir entscheiden, ob wir eine partielle Übersetzung vornehmen oder nicht, basierend auf den Punkten der Aufmerksamkeit.", "score": 70.0}533{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_234.wav", "doc_id": "oYCKgTzTDy.seg_234", "src_text": "We also test Monolingual Few-shot setting by training monolingual models with only 10% of training data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir testen auch die monolingualen Einstellungen, indem wir mit nur zwölf Prozent der Trainingsdaten monolinguale Modelle trainieren.", "score": 65.0}534{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_523.wav", "doc_id": "dvGkKzmIaN.seg_523", "src_text": "The trigger set is a group of words in a moderate frequency interval.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "die Antriebsmenge ist eine Gruppe von Wörtern in einem moderaten Frequenzintervall.", "score": 72.0}535{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_359.wav", "doc_id": "gGbuDbHhyc.seg_359", "src_text": "First, we find that, interestingly, recent WSL methods indeed require clean validation samples to work properly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wir fest, dass interessante neuere WSL-Methoden tatsächlich saubere Validierungsmuster erfordern, um richtig zu funktionieren.", "score": 94.0}536{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_510.wav", "doc_id": "dvGkKzmIaN.seg_510", "src_text": "To protect the copyright of embedding as services, one of the solutions is to embed a watermark in the provider service and detect whether another service contain the watermark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Um das Urheberrecht von eingebetteten Diensten zu schützen, wird eine der Lösungen ein Wasserzeichen in den Dienst des Anbieters eingebettet und festgestellt, ob ein anderes Dienst das Wasserzeichen enthält.", "score": 60.0}537{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_248.wav", "doc_id": "oYCKgTzTDy.seg_248", "src_text": "I think this is known as the \"Curse of Multilinguality\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "erzielt. Ich glaube, das ist ein Fluch der Vielsprachigkeit.", "score": 91.0}538{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_772.wav", "doc_id": "WTTtiRKFZI.seg_772", "src_text": "As you may know, there are different dependency structures assumed by different theories and corpus approaches.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wie Sie vielleicht wissen, werden verschiedene Abhängigkeitsstrukturen von verschiedenen Theorien und Korpusansätzen angenommen, also", "score": 97.0}539{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_235.wav", "doc_id": "oYCKgTzTDy.seg_235", "src_text": "And we test Multilingual Model which we train one multilingual model for all languages.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und es hat ein monolinguales Modell, das wir für alle Sprachen trainieren.", "score": 92.0}540{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_43.wav", "doc_id": "aQpIWggfCo.seg_43", "src_text": "We use large language models to generate a high-quality script dataset, CoScript, for constrained language planning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wen wenn nicht?", "score": 41.0}541{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_95.wav", "doc_id": "uZBWfYjYnf.seg_95", "src_text": "And what are the problems of the current SimulST models?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Und was sind die Probleme der aktuellen SimulST-Modelle?", "score": 94.0}542{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_189.wav", "doc_id": "SLpqvupgvW.seg_189", "src_text": "When we move higher in the list, the entities become more similar to each other and it's usually harder to make the disambiguation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir uns weiter oben auf der Liste bewegen, werden die Entitäten sich gegenseitig ähnlicher und es ist in der Regel schwieriger, die Abweichung zu erkennen.", "score": 88.0}543{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_751.wav", "doc_id": "XejEJmgUmE.seg_751", "src_text": "Finally, we can choose sentences from a completely unrelated domain such as Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Schließlich können wir Sätze aus einem völlig unabhängigen Domäne, wie z.B. einer", "score": 69.0}544{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_709.wav", "doc_id": "oaOHnMCwad.seg_709", "src_text": "For example, we find that data sets and models are most aligned to English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Beispiel finden wir heraus, dass die Datensätze und Modelle am meisten mit englischsprachigen Ländern ausgerichtet sind.", "score": 90.0}545{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_155.wav", "doc_id": "wLqFAuDnKa.seg_155", "src_text": "However, the \"Style/Awkward\" category for PaLM is lower than for the state-of-the-art systems, which is an additional signal that PaLM provides really fluent output, but still with some problems of accuracy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Allerdings ist die Kategorie „Stil“ für Panm niedriger als für die neuesten Systeme, was ein zusätzlicher Hinweis ist. Doch die Ausgabe ist wirklich fließend, aber mit einigen Problemen der Genauigkeit. Das", "score": 60.0}546{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_689.wav", "doc_id": "oaOHnMCwad.seg_689", "src_text": "So prior work has suggested some anecdotal evidence of having positionality, such as cultural gaps and models and data sets, as well as theoretical definitions of model positionality.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "So legen Prinziparbeiter einige anekdotische Beweise für die Positionierung voraus, wie kulturelle Lücken und Modelle und Datensätze, sowie die tatsächliche Definition der Modellposition.", "score": 60.0}547{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_262.wav", "doc_id": "oYCKgTzTDy.seg_262", "src_text": "Thanks for listening.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dank fürs Zuhören.", "score": 90.0}548{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_645.wav", "doc_id": "FLkGnzVRew.seg_645", "src_text": "To the goal of creating a cognitive dissonance resource, we conducted a large scale annotation of dissonance relations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zum Zweck der Schaffung einer kognitiven Distanzressource haben wir eine große Anzahl von Distanzbeziehungen hergestellt.", "score": 98.0}549{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_113.wav", "doc_id": "uZBWfYjYnf.seg_113", "src_text": "But also we want that they are shifted on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "auf diesem Plot. Aber auch wir wollen, dass sie auf der linken Seite verschoben werden.", "score": 65.0}550{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_813.wav", "doc_id": "WTTtiRKFZI.seg_813", "src_text": "But what's novel in this paper is that we observed that this tendency only occurs when the governor is on the left or absent.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Was jedoch neu ist, ist, dass wir festgestellt haben, dass diese Tendenz nur auftritt, wenn der Gouverneur links ist oder abwesend ist. Richtig, also", "score": 90.0}551{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_467.wav", "doc_id": "SUkmfOTvGi.seg_467", "src_text": "We observe that models have been used in CoNLL-2003 to develop NER for almost 20 years and this naturally raises several problems.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir beobachten, dass Modelle fast 20 Jahre lang Kondor verwenden, um NER zu entwickeln, und das bringt natürlich mehrere Probleme mit sich, zum Beispiel:", "score": 85.0}552{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_626.wav", "doc_id": "oeooqChmKK.seg_626", "src_text": "Additional experiments with fictional knowledge indicated even the best performing models, cannot reliably integrate backward knowledge provided only at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zusätzliche Experimente mit fiktivem Wissen zeigen, dass selbst die besten Modelle das Hintergrundwissen nicht zuverlässig integrieren können.", "score": 64.0}553{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_847.wav", "doc_id": "GvEBWkLmuI.seg_847", "src_text": "The benefit of this is that we get really specific stereotypes and patterns, without having to rely on any specific lexicon.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Der Vorteil davon ist, dass wir wirklich spezifische Stereotypen und Muster erhalten, ohne uns auf einen spezifischen Lexikon verlassen zu müssen.", "score": 58.0}554{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_797.wav", "doc_id": "WTTtiRKFZI.seg_797", "src_text": "It's okay the way instead of \"it\", we have this long NP.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist absolut faszinierend, ich bin okay, anstatt davon, dass wir die lange und lange Pinguine haben.", "score": 10.0}555{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_468.wav", "doc_id": "SUkmfOTvGi.seg_468", "src_text": "Firstly, can these models generalise to modern data?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Erstens, können diese Modelle auf moderne Daten generalisiert werden?", "score": 98.0}556{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_665.wav", "doc_id": "FLkGnzVRew.seg_665", "src_text": "On further rounds of AL with two best strategies, we improve dissonance classification AUC to 0.75, which is the best performance that we have on the task so far.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In weiteren Runden von AL mit zwei der besten Strategien verbesserten wir die Distanzklasse AUC auf 0,75, was bis jetzt die beste Leistung ist, die wir auf der Aufgabe erreicht haben.", "score": 95.0}557{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_264.wav", "doc_id": "PIZEXUFLAR.seg_264", "src_text": "So with the advances in large language models, many works started to explore new learning paradigms of reusing pre-trained language models for different downstream tasks in a parameter and data-efficient way.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Mit den Fortschritten bei großen Sprachmodellen begannen viele Arbeiten, neue Lernparadigmen für die Wiederverwendung von vortrainierten Sprachmodellen für verschiedene Downstream-Aufgaben zu erforschen. Viele Studien", "score": 57.0}558{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_443.wav", "doc_id": "hgIDlKNiFM.seg_443", "src_text": "To answer this question, we compare DrBERT with our ChuBERT model, which is based on anonymized data obtained from the Nantes University Hospital data warehouse.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Um diese Frage zu beantworten, vergleichen wir Dr. Bert mit unserem Schulbert-Modell, das auf anonymisierten Daten basiert, die wir aus dem Non University Hospital Data House erhalten.", "score": 95.0}559{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_59.wav", "doc_id": "TVCREhgqUP.seg_59", "src_text": "In particular, they often fail to reproduce the systematic correspondences between input and output, such as those that are color-coded in the example.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Insbesondere scheiterten sie oft daran, die systematischen Korrespondenzen zwischen Input und Output zu reproduzieren, wie die, die im Beispiel farbcodiert sind.", "score": 100.0}560{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_268.wav", "doc_id": "PIZEXUFLAR.seg_268", "src_text": "Additionally, at the time of our research, we discovered a considerable discrepancy in the availability of instructional datasets between NLP and multi-modal.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Darüber hinaus entdeckten wir in der Zeit unserer Forschung eine beträchtliche Diskrepanz in der Verfügbarkeit von Trainingsdatensätzen zwischen LBP und Multimodal.", "score": 97.0}561{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_860.wav", "doc_id": "GvEBWkLmuI.seg_860", "src_text": "So instead to do that, we'll turn to the results from our Marked Words method to show how these positive-seeming words facilitate stereotypes and essentializing narratives.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Stattdessen werden wir uns die Ergebnisse aus der Marktwahl ansehen, um zu zeigen, wie diese positiv erscheinenden Wörter Stereotypen und Stereotypisierungen aufweisen.", "score": 94.0}562{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_81.wav", "doc_id": "TVCREhgqUP.seg_81", "src_text": "Our model outperforms the others by a large margin on generalization to deeper recursion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unser Modell übertrifft die anderen bei der Generalisierung zu tiefer Rekursion deutlich. Andere", "score": 70.0}563{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_696.wav", "doc_id": "oaOHnMCwad.seg_696", "src_text": "And so we opt to re annotate data to get many annotates for instance and to get a rich set of demographic data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "werden, und so versuchen wir, Daten zu wiederauswerten, um viele Annotatoren für jede Instanz zu erhalten und zu teilen. -Set anhand der demografischen Daten.", "score": 100.0}564{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_277.wav", "doc_id": "PIZEXUFLAR.seg_277", "src_text": "We follow the method from OFA and formulate all the tasks in a unified sequence-to-sequence format.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir folgen dem von OFA gegebenen Leitfaden und formulieren alle Aufgaben in einer sequenz-zu-sequenz-Formatierung,", "score": 100.0}565{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_642.wav", "doc_id": "FLkGnzVRew.seg_642", "src_text": "High cognitive dissonance is also related to anxiety disorders and can help understand people's mental health better.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Eine hohe kognitive Distanz ist auch mit Angststörungen verbunden und kann helfen, Menschen ihre geistige Gesundheit besser zu verstehen.", "score": 96.0}566{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_143.wav", "doc_id": "wLqFAuDnKa.seg_143", "src_text": "It's the examples that carry most of the weight.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "größten Teil des Gewichts haben. Die", "score": 9.0}567{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_776.wav", "doc_id": "WTTtiRKFZI.seg_776", "src_text": "So these two approaches are asymmetric.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "also diese", "score": 3.0}568{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_786.wav", "doc_id": "WTTtiRKFZI.seg_786", "src_text": "OK.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "das", "score": 0.0}569{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_311.wav", "doc_id": "dJGfOSFgZO.seg_311", "src_text": "This work was done by the Emory NLP Lab led by Professor Jinho Choi at Emory University and in collaboration with Amazon Alexa AI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Arbeit wurde vom Emery Lab der Universität von Emery geleitet, unter der Leitung von Professor Gino Ochoa und in Zusammenarbeit mit Amazon Alexa AI. Lassen", "score": 36.0}570{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_394.wav", "doc_id": "WBLMIsdIrq.seg_394", "src_text": "In this work, we try to answer these two questions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In dieser Arbeit versuchen wir, diese beiden Fragen zu beantworten:", "score": 94.0}571{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_471.wav", "doc_id": "SUkmfOTvGi.seg_471", "src_text": "To investigate these problems, we developed the CoNLL++ Dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Um diese Probleme zu untersuchen, entwickeln wir das Daten-Satz Carneal", "score": 98.0}572{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_16.wav", "doc_id": "aQpIWggfCo.seg_16", "src_text": "Then we conduct detailed analysis to investigate why learning models fail.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dann führen wir detaillierte Analysen durch, um zu ergründen, was die Landmodellfunktionen sind. Die Ergebnisse", "score": 68.0}573{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_27.wav", "doc_id": "aQpIWggfCo.seg_27", "src_text": "We only keep the script if the target goal scores the highest in the goal set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "wenn das Skript die höchste Punktzahl im Zielbereich hat.", "score": 67.0}574{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_513.wav", "doc_id": "dvGkKzmIaN.seg_513", "src_text": "Second, the watermark should not degrade the utility of the provided embeddings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zweitens sollte die Wasserzeichenmethode die Nützlichkeit der vorgesehenen Einbauten nicht verringern.", "score": 71.0}575{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_123.wav", "doc_id": "wLqFAuDnKa.seg_123", "src_text": "This is joint work with my colleagues from Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies ist eine gemeinsame Arbeit mit meinen Kollegen von Google Translate.", "score": 100.0}576{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_594.wav", "doc_id": "oeooqChmKK.seg_594", "src_text": "For example, in the sentence, \"John saw the newly elected president on TV.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Beispiel sah John in der Sätze den neu gewählten Präsidenten im Fernsehen.", "score": 60.0}577{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_824.wav", "doc_id": "WTTtiRKFZI.seg_824", "src_text": "And we show in the paper how this provides an argument against asymmetric structures of coordination, as these two, and for the symmetric structures, as these two.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und wir zeigen im Papier, wie dies geschieht. 1 bietet ein Argument gegen asymmetrische Koordinierungsstrukturen wie diese beiden und fördert asymmetrische Strukturen wie diese beiden. Sehen", "score": 60.0}578{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_107.wav", "doc_id": "uZBWfYjYnf.seg_107", "src_text": "For example, if we receive a speech chunk containing \"I'm going to talk about...\" and our model predicts the translation in German, and we will look at the cross-attention weights, we'll see that the first two words points to the earliest received speech frames, while the last word points to the last received speech frames, as lambda speech frames.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir beispielsweise einen Spruchabschnitt erhalten, der „Ich werde darüber sprechen“ enthält, und unser Modell eine Übersetzung ins Deutsche vorhersagt, wird die Übersetzung in der Regel verwendet. Und wir werden uns die Gewichte ansehen. Wir können sehen, dass die ersten beiden Wörter auf die frühesten erhaltenen Sprachrahmen hinweisen, während das letzte Wort auf die letzten erhaltenen Sprachrahmen hinweist („Lambda-Sprachrahmen“).", "score": 60.0}579{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_3.wav", "doc_id": "aQpIWggfCo.seg_3", "src_text": "Previous work has exploited language models to plan for abstract goals of stereotypical activities such as \"make a cake\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "vorherigen Arbeit wurden Sprachmodelle genutzt, um für abstrakte Ziele von stereotypischen Aktivitäten wie Make-a-Kick zu planen und zu", "score": 60.0}580{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_431.wav", "doc_id": "hgIDlKNiFM.seg_431", "src_text": "Hi, I am Yanis Labrak and I will present you our works on \"DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical Domains.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hallo, ich bin Yanislav und werde Ihnen unsere Arbeiten in Französisch für biomedizinische und klinische Bereiche vorstellen.", "score": 60.0}581{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_753.wav", "doc_id": "XejEJmgUmE.seg_753", "src_text": "So how does the model do?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "uns ansehen. Also, wie funktioniert das Modell?", "score": 58.0}582{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_566.wav", "doc_id": "rISrKoXQCx.seg_566", "src_text": "So we divide pretraining corpora, into pre 45th president of the United States and after 45th president of the United States.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir unterteilen also das Präsidium der Vereinigten Staaten in zwei verschiedene temporale Korpora und das Präsidium der Vereinigten Staaten", "score": 52.0}583{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_101.wav", "doc_id": "uZBWfYjYnf.seg_101", "src_text": "First, to use already existing offline ST models without re-training or adopting specific architecture for SimulST.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zunächst verwenden wir bereits bestehende Offline-CT-Modelle ohne Wiedertrainieren oder spezifische Architekturen für CT-CT anpassen:", "score": 56.0}584{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_274.wav", "doc_id": "PIZEXUFLAR.seg_274", "src_text": "For investigating multi-modal instruction tuning on our proposed dataset, we take OFA, a unified multi-modal pre-trained model, as our base model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Für die Untersuchung der mehrmodalen Anweisung auf unserem vorgeschlagenen Datensatz nehmen wir Ofa als unser Basismodell, Ofa", "score": 70.0}585{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_495.wav", "doc_id": "SUkmfOTvGi.seg_495", "src_text": "So going back to the question that we raised in the title of our paper Do CoNLL-2003 taggers still work in 2023?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zurück zu der Frage, die wir in unserem Bericht aufgeworfen haben: Funktionieren die Cornel-Tagger noch im Jahr 2003? Und", "score": 36.0}586{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_301.wav", "doc_id": "PIZEXUFLAR.seg_301", "src_text": "As we can see by transfer learning from natural instruction datasets, the model can achieve much better sensitivity compared to the original OFA model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Transferlernen von den Datensätzen der natürlichen Anweisung kann das Modell im Vergleich zum ursprünglichen OA-Modell eine viel höhere Empfindlichkeit erreichen.", "score": 51.0}587{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_68.wav", "doc_id": "TVCREhgqUP.seg_68", "src_text": "Our approach predicts the output from the input in two steps.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Unser Ansatz prognostiziert den Output aus dem Input in zwei Schritten. Zuerst", "score": 84.0}588{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_723.wav", "doc_id": "oaOHnMCwad.seg_723", "src_text": "I mean, we want to emphasise that inclusive NLP isn't just making.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist die Masakani-Initiative. Ich möchte betonen, dass die inklusive NLP nicht nur alle", "score": 66.0}589{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_258.wav", "doc_id": "oYCKgTzTDy.seg_258", "src_text": "We conduct a comprehensive benchmark study on three representative types of multilingual language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir führen eine umfassende Benchmark-Studie zu drei Vertretern von Typen von Mehrsprachigen Modellen durch,", "score": 83.0}590{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_288.wav", "doc_id": "PIZEXUFLAR.seg_288", "src_text": "In each experiment, we report the min and max performance and the standard deviation of the performance across all 5 experiments.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "in jedem Experiment bewerten. Wir berichten über die mittlere und maximale Leistung und die Standardabweichung der Leistung in allen fünf Experimenten.", "score": 93.0}591{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_583.wav", "doc_id": "rISrKoXQCx.seg_583", "src_text": "If we do try to sanitaze somehow, we would also risk censorship, or exclusion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wenn wir versuchen würden, in irgendeiner Weise zu sanitieren, würden wir auch Zensur oder Auslassungen riskieren, und", "score": 81.0}592{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_179.wav", "doc_id": "SLpqvupgvW.seg_179", "src_text": "In the second speech bubble, Alice says, \"Do you mean 'Easy on Me' or 'I Gotta Feeling'?\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In der zweiten Sprechblase sagt Alice: „Meinst du leicht von mir oder habe ich ein Gefühl?“", "score": 85.0}593{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_829.wav", "doc_id": "GvEBWkLmuI.seg_829", "src_text": "This work is done in collaboration with Esin Durmus and Dan Jurafsky.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "in Zusammenarbeit mit Esender und Danroski durchgeführt.", "score": 63.0}594{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_564.wav", "doc_id": "rISrKoXQCx.seg_564", "src_text": "For example, for RoBERTa further trained on the left-leaning Reddit corpus we can see a substantial liberal shift in terms of its political biases.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Beispiel, für Roberta, weiter trainiert auf dem linken linkierten Korpus, können wir einen substantiellen liberalen Verschiebung in In Bezug auf die politischen Vorurteile", "score": 66.0}595{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_198.wav", "doc_id": "SLpqvupgvW.seg_198", "src_text": "Here's for example, the Google search result for the song \"Easy on Me.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hier ist zum Beispiel das Google-Suchergebnis für das Lied „Easy Annie“. Für", "score": 66.0}596{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_778.wav", "doc_id": "WTTtiRKFZI.seg_778", "src_text": "They single out one of the conjuncts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Ansätze sind symmetrisch. Jetzt sind auch die", "score": 0.0}597{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_498.wav", "doc_id": "SUkmfOTvGi.seg_498", "src_text": "And lastly, please make sure to check out our paper, our data set and if you have any questions, feel free to contact me.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "bitte, überprüfen Sie unser Papier, unsere Datenbank, und wenn Sie irgendwelche Fragen haben, können Sie sich frei mit mir in Verbindung setzen.", "score": 90.0}598{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_260.wav", "doc_id": "oYCKgTzTDy.seg_260", "src_text": "And et cetera.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "usw. und wir", "score": 95.0}599{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_26.wav", "doc_id": "aQpIWggfCo.seg_26", "src_text": "In addition, we reward the script that contains the keywords of the target constraint.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Darüber hinaus achten wir auf das Skript, das die Schlüsselwörter des Zielkontrahnts enthält.", "score": 80.0}600{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_821.wav", "doc_id": "WTTtiRKFZI.seg_821", "src_text": "So I'll concentrate on the right one.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "ermittelt werden kann. Was wir sagen, ist, dass die Regierung auf der", "score": 0.0}601{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_469.wav", "doc_id": "SUkmfOTvGi.seg_469", "src_text": "And when we develop new taggers, what is needed for good generalization?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und wenn wir neue Tags entwickeln, was ist für eine gute Generalisierung erforderlich?", "score": 91.0}602{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_522.wav", "doc_id": "dvGkKzmIaN.seg_522", "src_text": "Before these main steps, we first select a trigger set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hauptschritte durchgehen, wählen wir zunächst ein Triggerset.", "score": 90.0}603{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_674.wav", "doc_id": "oaOHnMCwad.seg_674", "src_text": "Hi everyone.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo,", "score": 89.0}604{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_752.wav", "doc_id": "XejEJmgUmE.seg_752", "src_text": "So this will tell us like whether the models acceptability judgments are actually impacted by any context, like, whether the context is coming from a different subset of the data set, or whether it's like completely irrelevant, to the current like to the sentence that we are looking at.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "anderen Sprache oder einer anderen Domäne, auswählen, um die Akzeptanzfähigkeit der Modelle zu testen. Wikipedia, auswählen. Erzählen Sie uns, ob die Akzeptabilitätsurteile der Modelle tatsächlich von irgendeinem Kontext beeinflusst werden, wie zum Beispiel, ob der Kontext aus einem anderen Teilmenge des Datensatzes kommt oder ob er völlig irrelevant zum aktuellen Satz ist, den wir", "score": 61.0}605{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_109.wav", "doc_id": "uZBWfYjYnf.seg_109", "src_text": "If we go on and we receive another speech chunk, and our model predicts other three words and we will look at those cross-attention weights, we will see that no word points to the last lambda speech frames.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn wir weitermachen und wir erhalten einen anderen Speech-Tag und unser Modell prädiziert. Wir werden die drei Wörter in der Reihenfolge, in der sie geordnet sind, und wir werden die Kreuz-Atten-Wege darauf untersuchen. Wir werden sehen, dass kein Wort auf die letzten Lambdas des Speech-Frames zeigt.", "score": 63.0}606{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_721.wav", "doc_id": "oaOHnMCwad.seg_721", "src_text": "Our third recommendation is to build specialised datasets and models within 4 specific communities.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "dritte Empfehlung ist es, spezielle Datensätze und Modelle in vier spezifischen Gemeinschaften zu erstellen,", "score": 82.0}607{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_332.wav", "doc_id": "dJGfOSFgZO.seg_332", "src_text": "These reliable, informative, and distinct ABC-Eval metrics enable us to evaluate conversational AI with a higher resolution than previous methods are able to achieve.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Diese zuverlässigen, informierenden und einzigartigen ABC-EVL-Metriken ermöglichen es uns, die konversationelle AI mit einer höheren Auflösung zu bewerten als es die vorherigen Methoden erreichen können.", "score": 88.0}608{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_238.wav", "doc_id": "oYCKgTzTDy.seg_238", "src_text": "And we also consider Cross-lingual Zero-shot and Few-shot transfer.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wir erwägen auch die Transferierung von Zero-Shot- und Feature-Transfer", "score": 74.0}609{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_564.wav", "doc_id": "rISrKoXQCx.seg_564", "src_text": "For example, for RoBERTa further trained on the left-leaning Reddit corpus we can see a substantial liberal shift in terms of its political biases.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Beispielsweise können wir bei Roberta eine weitergehende Finanzierung des linken Korpus sehen. In Bezug auf seine politischen Vorurteile.", "score": 60.0}610{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_87.wav", "doc_id": "TVCREhgqUP.seg_87", "src_text": "We address this by inducing the alignment as part of the training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir behandeln dies, indem wir die Ausrichtung als Teil des Trainings induzieren.", "score": 94.0}611{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_15.wav", "doc_id": "aQpIWggfCo.seg_15", "src_text": "We find that all language models achieve unsatisfactory results on planning for specific goals.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir stellen fest, dass alle linearen Modelle unzufriedenstellende Ergebnisse bei der Planung für spezifische Ziele erzielen.", "score": 53.0}612{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_9.wav", "doc_id": "aQpIWggfCo.seg_9", "src_text": "A good planner should write scripts that are reasonable and faithful to constraints.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ein guter Planer sollte Skripte schreiben, die vernünftig und den Einschränkungen treu sind.", "score": 88.0}613{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_52.wav", "doc_id": "TVCREhgqUP.seg_52", "src_text": "As usual, we have a training set of utterances.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "es so aus, als hätten Sie in", "score": 0.0}614{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_112.wav", "doc_id": "uZBWfYjYnf.seg_112", "src_text": "So we want our curves to be as high as possible on this plot.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "dass unsere Queue so hoch wie möglich auf", "score": 80.0}615{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_455.wav", "doc_id": "hgIDlKNiFM.seg_455", "src_text": "We also observe that using more data translated to better performance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "um mehrsprachig zu sein, und dass die Verwendung mehrerer Daten zu besseren Leistungen führt.", "score": 0.0}616{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_371.wav", "doc_id": "gGbuDbHhyc.seg_371", "src_text": "So in practice, there's no reason to choose more complex WSL methods which require more computation time and disk space.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daher gibt es keinen Grund, in der Praxis komplexere WS-L-Methoden zu wählen, die mehr Berechnungszeit und Diskraum benötigen.", "score": 100.0}617{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_659.wav", "doc_id": "FLkGnzVRew.seg_659", "src_text": "\"Cumulative\" accumulates all the data collected from active annotation so far, whereas \"Iterative\" updates the model by training on the latest set of data collected.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Alle Daten aus den aktiven Lern- und Anmerkungsrunden werden kumuliert, um das Modell schrittweise zu aktualisieren. Bei den", "score": 10.0}618{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_348.wav", "doc_id": "gGbuDbHhyc.seg_348", "src_text": "If we directly train neural networks on weakly labeled data, the neural networks tend to memorize the label noise and do not generalize.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir neuronale Netze direkt trainieren und schwach beschriftete Daten haben, tendieren die neuronalen Netze dazu, den Label-Rauschen zu memorieren und nicht zu", "score": 92.0}619{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_244.wav", "doc_id": "oYCKgTzTDy.seg_244", "src_text": "We found that Encoder-Decoder obtains the best performance on all nine datasets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir stellten fest, dass der Encoder-Decoder die beste Leistung auf allen neun Datensätzen erzielt.", "score": 100.0}620{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_136.wav", "doc_id": "wLqFAuDnKa.seg_136", "src_text": "And this can go, in extreme cases, up to 40 BLEURT points.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Und dies kann in extremen Fällen bis zu 40 Punkte", "score": 68.0}621{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_774.wav", "doc_id": "WTTtiRKFZI.seg_774", "src_text": "So in this case, Lisa.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist, in diesem Fall Lisa.", "score": 92.0}622{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_744.wav", "doc_id": "XejEJmgUmE.seg_744", "src_text": "And what we do is that to recreate like longer sequences and which are acceptable and which has the same matching of the grammatical structure.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und was wir tun, um längere Sequenzen zu erzeugen, die akzeptabel sind und die gleiche grammatikalische Struktur haben,", "score": 93.0}623{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_279.wav", "doc_id": "PIZEXUFLAR.seg_279", "src_text": "Ok, now I'm going to talk about multi-modal instruction tuning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Ok, ich werde jetzt über die Multi-Modell-Unterstützung sprechen.", "score": 100.0}624{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_324.wav", "doc_id": "dJGfOSFgZO.seg_324", "src_text": "For comparison, we also evaluated these conversations using three existing methods: Likert ratings on the turn-level, Likert ratings on the dialogue-level, and dialogue-level pairwise comparisons.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zum Vergleich bewerteten wir diese Gespräche mit drei existierenden Methoden: Likert-Skalen auf der Wendebene, Likert-Skalen auf der Dialogebene und Paarweisen-Vergleiche auf der Dialogebene.", "score": 99.0}625{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_192.wav", "doc_id": "SLpqvupgvW.seg_192", "src_text": "The third one is when they have similar descriptions on Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Der dritte ist, wenn sie ähnliche Beschreibungen auf Wikipedia haben", "score": 99.0}626{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_212.wav", "doc_id": "SLpqvupgvW.seg_212", "src_text": "We've also shown that the models are domain-generalizable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir zeigen auch, dass die Modelle domänenübergreifend sind.", "score": 100.0}627{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_402.wav", "doc_id": "WBLMIsdIrq.seg_402", "src_text": "Now we analyze words with high P-CXMI to look for patterns between these words.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Jetzt analysieren wir Wörter mit hohen PSMI, um nach Übereinstimmungen zwischen diesen Wörtern zu suchen.", "score": 37.0}628{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_381.wav", "doc_id": "gGbuDbHhyc.seg_381", "src_text": "Please feel free to check it out.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "finden können. Bitte fühlen Sie sich frei, ihn zu überprüfen.", "score": 58.0}629{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_107.wav", "doc_id": "uZBWfYjYnf.seg_107", "src_text": "For example, if we receive a speech chunk containing \"I'm going to talk about...\" and our model predicts the translation in German, and we will look at the cross-attention weights, we'll see that the first two words points to the earliest received speech frames, while the last word points to the last received speech frames, as lambda speech frames.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Beispielsweise, wenn wir einen Sprachblock mit der Übersetzung „Ich werde darüber sprechen“ erhalten und unser Modell eine Übersetzung in Deutsch vorhersagt. Und wir werden auf die Querverbindung achten. Wir werden sehen, dass die ersten beiden Wörter auf die frühesten erhaltenen Sprachrahmen verweisen, während das letzte Wort auf die letzten erhaltenen Sprachrahmen verweist.", "score": 96.0}630{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_727.wav", "doc_id": "oaOHnMCwad.seg_727", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 100.0}631{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_456.wav", "doc_id": "hgIDlKNiFM.seg_456", "src_text": "Overall, from-scratch pre-training seems to obtain higher performance on most of the tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Insgesamt scheint das Training mit Schrammen eine höhere Leistung auf den meisten Aufgaben zu erzielen. Unsere Experimente mit kontinuierlicher", "score": 63.0}632{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_625.wav", "doc_id": "oeooqChmKK.seg_625", "src_text": "This suggests that when trained on generic reference resolution data sets, most learn to exploit surface cues, which are not useful when testing on KITMUS where such queues have been removed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies deutet darauf hin, dass trainierte Mäuse lernen, Oberflächenhinweise auszunutzen, die bei der Untersuchung von Kitts, wo solche Hinweise entfernt wurden, nicht nützlich sind.", "score": 44.0}633{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_531.wav", "doc_id": "dvGkKzmIaN.seg_531", "src_text": "We first construct a back door and a benign data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir bauen zunächst eine Rückwand und einen bösartigen", "score": 49.0}634{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_117.wav", "doc_id": "uZBWfYjYnf.seg_117", "src_text": "And we see that it outperforms all the strategies applied to offline models since the curves are shifted over the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir sehen, dass EAD alle Strategien, die auf Offline-Modelle angewendet werden, übertrifft, da ihre Kurven nach links verschoben sind.", "score": 90.0}635{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_638.wav", "doc_id": "FLkGnzVRew.seg_638", "src_text": "And they have a consonance relationship.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und sie haben eine konsensuale Beziehung. Die Diskrepanz", "score": 43.0}636{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_311.wav", "doc_id": "dJGfOSFgZO.seg_311", "src_text": "This work was done by the Emory NLP Lab led by Professor Jinho Choi at Emory University and in collaboration with Amazon Alexa AI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Diese Arbeit wurde von dem Emory NLP-Labor geleitet von Professor Jino Choi an der Emory University und in Zusammenarbeit mit Amazon Alexa erstellt. Lassen", "score": 73.0}637{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_321.wav", "doc_id": "dJGfOSFgZO.seg_321", "src_text": "ABC-Eval is capable of measuring the rates at which chat models will commit various thematic errors.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ABC-EVAL kann die Raten messen, mit denen Chat-Modelle verschiedene thematische Fehler begehen.", "score": 100.0}638{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_191.wav", "doc_id": "SLpqvupgvW.seg_191", "src_text": "The second one is when the entities have similar titles, for example, two books with the name \"The Return\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das zweite ist, wenn die Einheiten ähnliche Titel haben, z. B. zwei Bücher mit dem Namen „The Rite“.", "score": 65.0}639{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_207.wav", "doc_id": "SLpqvupgvW.seg_207", "src_text": "If the language model has access to the exact same background knowledge as the annotators, then the accuracy is really high, it's around 92 to 95%.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn das Sprachmodell Zugriff auf die exakt gleiche Hintergrundwissensbasis wie die Annotatoren hat, ist die Genauigkeit wirklich hoch: sie liegt bei etwa 92-95. Aber", "score": 96.0}640{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_551.wav", "doc_id": "rISrKoXQCx.seg_551", "src_text": "This has created a mixed blessing for language model applications.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Lobeshymne für eine Sprachmodellanwendung geschaffen. So können sie auf", "score": 68.0}641{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_236.wav", "doc_id": "oYCKgTzTDy.seg_236", "src_text": "For example, we put the German, English, Chinese queries together to train a multilingual model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zum Beispiel vereinigen wir die deutschen, englischen und chinesischen Fragen, um ein mehrsprachiges Modell zu trainieren,", "score": 100.0}642{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_737.wav", "doc_id": "XejEJmgUmE.seg_737", "src_text": "The current MPP pipeline basically doesn't allow us to evaluate a model's acceptance towards longer sentences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die aktuelle MP-Pipeline erlaubt es uns im Grunde nicht, die Akzeptanz eines Modells für längere Sätze zu bewerten.", "score": 100.0}643{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_51.wav", "doc_id": "TVCREhgqUP.seg_51", "src_text": "In the context of semantic parsing, testing for compositional generalization might look like this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Im Kontext der semantischen Parsenierung, die für die kompositorische Generalisierung getestet wird, könnte", "score": 79.0}644{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_856.wav", "doc_id": "GvEBWkLmuI.seg_856", "src_text": "However, when we actually look at the distribution of the words and lexicon, we find very different things.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir uns jedoch die Verteilung der Wörter in einem Wörterbuch anschauen, stellen wir jedoch fest, dass es sehr unterschiedliche Dinge", "score": 85.0}645{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_478.wav", "doc_id": "SUkmfOTvGi.seg_478", "src_text": "The first one is the model architecture.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die erste ist die Modellarchitektur.", "score": 100.0}646{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_245.wav", "doc_id": "oYCKgTzTDy.seg_245", "src_text": "And we evaluate on mT5 and XLM-R + PTR on multilingual setting.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und wir bewerten auf M.T.5 und als Beispiel XLM-R + PDR auf einer multilinguellen Einstellung.", "score": 100.0}647{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_680.wav", "doc_id": "oaOHnMCwad.seg_680", "src_text": "But that's not really the case for Aditya Sharma.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Aber das ist wirklich der Fall für Aditya Sharma, die", "score": 72.0}648{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_521.wav", "doc_id": "dvGkKzmIaN.seg_521", "src_text": "Watermark injection and copyright verification.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wasserzeichen-Injektion und eine Urheberrechtsverwertung.", "score": 100.0}649{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_718.wav", "doc_id": "oaOHnMCwad.seg_718", "src_text": "So we have a few recommendations for this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Also haben wir einige Empfehlungen", "score": 97.0}650{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_641.wav", "doc_id": "FLkGnzVRew.seg_641", "src_text": "Studying cognitive dissonance can help us understand the effects of disagreement among people, track trends and belief values, and attitude changes in population.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das Studium der konzeptuellen Distanz kann helfen, die Auswirkungen von Meinungsverschiedenheiten unter Menschen, Trends und Überzeugungen, Werten und Einstellungen in der Bevölkerung zu verstehen.", "score": 57.0}651{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_639.wav", "doc_id": "FLkGnzVRew.seg_639", "src_text": "While dissonance is a very common phenomenon we experienced in daily decision making, they are really rare to find expressed in language among other kinds of discourse relations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist ein sehr häufiges Phänomen, das wir in der täglichen Entscheidungsfindung erleben, und sie ist wirklich bereit, in einer Sprache ausdrücklich auszudrücken, die wir in anderen Diskussionen nicht verwenden.", "score": 60.0}652{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_522.wav", "doc_id": "dvGkKzmIaN.seg_522", "src_text": "Before these main steps, we first select a trigger set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vor diesen Hauptschritten wählen wir zunächst eine Antriebsmenge aus;", "score": 84.0}653{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_470.wav", "doc_id": "SUkmfOTvGi.seg_470", "src_text": "At the same time, if we do observe poor generalization, what causes the performance drop of these models?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Gleichzeitig, wenn wir eine schlechte Generalisierung beobachten, was verursacht die Leistungseinbußen dieser Modelle?", "score": 93.0}654{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_259.wav", "doc_id": "oYCKgTzTDy.seg_259", "src_text": "And our results show many interesting findings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und unsere Ergebnisse zeigen viele interessante Erkenntnisse,", "score": 100.0}655{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_654.wav", "doc_id": "FLkGnzVRew.seg_654", "src_text": "We transfer from two different tasks: topic independent dissonance stance classification, a task that determines if two debate statements from different people are in agreement or in disagreement, irrespective of topic, called debate here, and on binary classification of expansion and comparison classes of PDTB since these two are closely related to the conception of consonance and dissonance and we call them CE here.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir übertragen von zwei verschiedenen Themen: Thema unabhängig von der Standortklassifizierung, die bestimmt, ob zwei Aussagen von verschiedenen Personen in Übereinstimmung oder in Unmöglichkeit sind, unabhängig vom Thema. Hier wird eine Debatte geführt und über die binäre Klassifizierung von Expansion und Vergleichsklassen von PendetB gesprochen, da diese beiden eng mit dem Konzept von Konsonanten und Dissonanzen verwandt sind, und wir nennen sie hier CeE.", "score": 15.0}656{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_857.wav", "doc_id": "GvEBWkLmuI.seg_857", "src_text": "So, while the generated personas have much higher rates of the lexicon words, the human-written ones have a much wider distribution of words, while the stereotype words that are in the generated personas are really just the words \"tall\" and \"athletic\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "die generierten Personas viel höhere Raten der Lexikonwörter haben. Die von Menschen geschriebenen Texte haben eine viel breitere Verteilung von Wörtern, während die stereotypen Wörter, die in den generierten Persönlichkeiten vorkommen, wirklich nur die Wörter \"toll\" und \"athletisch\" sind.", "score": 79.0}657{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_830.wav", "doc_id": "GvEBWkLmuI.seg_830", "src_text": "In recent years, many have documented the prevalence of social bias and stereotypes in large language models, or LLMs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In den letzten Jahren haben viele die Vorherrschaft des sozialen Biaßes und Stereotypen in großen Sprachmodellen oder LLMs dokumentiert.", "score": 100.0}658{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_823.wav", "doc_id": "WTTtiRKFZI.seg_823", "src_text": "But when the governor is on the right this tendency disappears.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "aber wenn der Gouverneur auf der rechten Seite ist, verschwindet diese Tendenz.", "score": 100.0}659{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_602.wav", "doc_id": "oeooqChmKK.seg_602", "src_text": "Kea is a Baker.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Kiah ist Bäcker.", "score": 99.0}660{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_350.wav", "doc_id": "gGbuDbHhyc.seg_350", "src_text": "In recent works in WSL, so WSL stands for Weakly Supervised Learning, a common claim is that people say that they only train models on the weakly labeled data and achieve high performance on clean test sets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "WSL ein Akronym für wöchentliches Superwise-Lernen. Eine häufige Behauptung ist, dass Menschen nur Modelle unter wöchentlichem Label-Data trainieren und auf sauberen Test-Sets hohe Leistung erzielen.", "score": 69.0}661{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_492.wav", "doc_id": "SUkmfOTvGi.seg_492", "src_text": "Our conclusion is that, for good generalization we would need a better model architecture, larger model size, as well as more fine tuning examples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unser Schluss ist, dass wir für eine gute Generalisierung eine bessere Modellarchitektur, eine größere Modellgröße sowie mehr feinjustierte Beispiele", "score": 98.0}662{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_217.wav", "doc_id": "oYCKgTzTDy.seg_217", "src_text": "So, semantic parsing is a task to build semantic representations of user queries such as SQL and Lambda Calculus.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "semantische Parsierung ist also eine Aufgabe, um semantische Darstellungen von Benutzeranfragen wie Zequel und Lambda-Kalküle zu bauen.", "score": 75.0}663{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_486.wav", "doc_id": "SUkmfOTvGi.seg_486", "src_text": "The second hypothesis is temporal drift which is the performance degradation that is caused by the increasing temporal gap between the train and the test data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die zweite Hypothese ist der temporale Abfall, der durch den zunehmenden zeitlichen Abstand zwischen Zug und Testdaten verursacht wird.", "score": 90.0}664{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_460.wav", "doc_id": "hgIDlKNiFM.seg_460", "src_text": "We are also observing that more specialized data is better, but it doesn't scale well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "dass spezialisierte Daten besser sind, mehr spezialisierte Daten besser sind, aber sie skaliert nicht gut, da", "score": 40.0}665{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_303.wav", "doc_id": "PIZEXUFLAR.seg_303", "src_text": "So overall, we propose the first large scale multi-model instruction tuning dataset with significantly improved their short capability of OFA, and we explore different transfer learning technique and show their benefits.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Insgesamt schlagen wir also ein erstes groß angelegtes multimodales Anweisungsjustierung-Datensatz vor, wir verbessern die Echtzeitfähigkeit von OFA erheblich und wir untersuchen verschiedene Transferlernmethoden und zeigen ihre Vorteile.", "score": 100.0}666{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_194.wav", "doc_id": "SLpqvupgvW.seg_194", "src_text": "For example, the same genre or the same artist for a song.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "z. B. das gleiche Genre oder der gleiche Künstler.", "score": 90.0}667{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_669.wav", "doc_id": "FLkGnzVRew.seg_669", "src_text": "In summary, we find that PRC is a simple AL strategy for rare class acquisition and cold starting AL with appropriately designed transfer learning task and help significantly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "stellen wir fest, dass die PRRC eine einfache AL-Strategie für die Aufnahme in die nächste Klasse ist und hilfreich ist, wenn man sie richtig anwendet.", "score": 50.0}668{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_799.wav", "doc_id": "WTTtiRKFZI.seg_799", "src_text": "So the reasoning here is that this is possible because even though this sentence violates the general grammatical principle that direct objects should be next to the verb, it satisfies the principle of dependency length minimization, which says that shorter dependencies are preferred.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Bienen, gelesen habe, also der Grund hierfür. Ist das möglich, weil, auch wenn diese Sätze die allgemeine grammatische Regel verletzen, dass ein direktes Objekt direkt nach dem Verb stehen sollte, sie den Prinzip der Abhängigkeitslänge minimierung erfüllen, das besagt, dass kürzere Abhängigkeiten bevorzugt werden.", "score": 60.0}669{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_438.wav", "doc_id": "hgIDlKNiFM.seg_438", "src_text": "Since its release in 2018, BERT has become one of the most effective approach to solve natural language processing tasks and offers huge performance gains compared to historical static and contextualized methods such as Word2vec, fastText, or more.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Seit seiner Veröffentlichung im Jahr 2018 ist BERT zu einem der effektivsten Ansätze zur Lösung von Aufgaben der natürlichen Sprachverarbeitung geworden und bietet einen enormen Leistungszuwachs im Vergleich zu historischen statischen und kontextualisierten Methoden wie Word2Vec oder GloVe. \"oder was?\".", "score": 85.0}670{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_536.wav", "doc_id": "dvGkKzmIaN.seg_536", "src_text": "Meanwhile, we also apply KS test and use its p-value as the third metric.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In der Zwischenzeit wenden wir auch den KS-Test an und verwenden sein p-Wert als dritte Metrik.", "score": 99.0}671{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_496.wav", "doc_id": "SUkmfOTvGi.seg_496", "src_text": "And we found that the answer is actually a resounding yes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "noch im Jahr 2023? Die Antwort ist tatsächlich ein eindeutiger „Ja“.", "score": 0.0}672{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_465.wav", "doc_id": "SUkmfOTvGi.seg_465", "src_text": "Let's get started.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "beginnen.", "score": 60.0}673{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_224.wav", "doc_id": "oYCKgTzTDy.seg_224", "src_text": "For example, there's only one single model to evaluate them.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zum Beispiel gibt es nur ein einziges Modell zur Bewertung. Zu", "score": 94.0}674{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_466.wav", "doc_id": "SUkmfOTvGi.seg_466", "src_text": "Our paper investigated the problem of generalization using the Named Entity Recognition Task or the NER task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In unserem Papier untersuchten wir das Problem der Generalisierung, indem wir die Aufgabe der Erkennung benannter Entitäten oder die NER-Aufgabe verwendeten.", "score": 100.0}675{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_392.wav", "doc_id": "WBLMIsdIrq.seg_392", "src_text": "Firstly because only a small portion of translations depend on context which makes corpus-level metrics like BLEU unable to capture these translations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "einmal, weil nur ein kleiner Teil der Übersetzungen von Kontext abhängt, was es Korpusniveau-Metriken wie Blue unmöglich macht, diese Übersetzungen zu erfassen. Und", "score": 99.0}676{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_180.wav", "doc_id": "SLpqvupgvW.seg_180", "src_text": "Which is the alternative question.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das ist die alternative", "score": 50.0}677{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_290.wav", "doc_id": "PIZEXUFLAR.seg_290", "src_text": "If it's a multi-modal generation task, we report Rouge-L. For NLP task, we report Rouge-L as well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn es sich um eine Multimodal-Generation handelt, berichten wir über die RuG. Für NRP-Aufgaben berichten wir auch über die RuG.", "score": 50.0}678{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_801.wav", "doc_id": "WTTtiRKFZI.seg_801", "src_text": "So here we have a dependency from \"read\" to the adjunct of length 7 measured in words and from \"read\" to \"book\" of length 4, so together it's 11.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hier haben wir also die Abhängigkeit von rot bis zum Adjektiv von Länge sieben gemessen in Wörtern und von rot bis zum Buch von Länge vier. Wenn Sie sich bewegen", "score": 40.0}679{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_734.wav", "doc_id": "XejEJmgUmE.seg_734", "src_text": "Which can also include grammaticality like BLiMP, SyntaxGym, or acceptability in terms of stereotypes such as CrowS pairs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die auch Grammatikalität wie „Blimp“ oder „Syntax Gem“ oder Akzeptabilität in Bezug auf Stereotypen wie „Crowds“ umfassen können.", "score": 100.0}680{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_427.wav", "doc_id": "WBLMIsdIrq.seg_427", "src_text": "We also compared different commercial systems and our benchmark shows that DeepL is usually more accurate than Google Translate for document-level translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir vergleichen auch unterschiedliche kommerzielle Systeme und unsere Benchmarks zeigen, dass die Google-Übersetzung für lokale Dokumentenübersetzung normalerweise genauer ist.", "score": 100.0}681{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_63.wav", "doc_id": "TVCREhgqUP.seg_63", "src_text": "This can be complicated and sometimes a computationally expensive process.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies kann ein komplizierter und manchmal rechenintensiver Prozess sein.", "score": 100.0}682{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_276.wav", "doc_id": "PIZEXUFLAR.seg_276", "src_text": "Here we show some example instances from our MultiInstruct dataset, to unify the processing of various input and output data types.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hier zeigen wir einige Beispielinstanzen aus unserem Multi-Instanz-Datensatz. Um die Verarbeitung verschiedener Eingabe- und Ausgabedatentypen zu vereinheitlichen,", "score": 66.0}683{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_615.wav", "doc_id": "oeooqChmKK.seg_615", "src_text": "This last setting is especially interesting, since it simulates the case where the background knowledge necessary to solve a task is not part of the pretrain data of models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die letzte Einstellung ist besonders interessant, da sie den Fall simuliert, bei dem das Hintergrundwissen zur Lösung einer Aufgabe nicht Teil der vorgefertigten Modelle ist, da sich", "score": 70.0}684{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_54.wav", "doc_id": "TVCREhgqUP.seg_54", "src_text": "And \"Mary knew that the girl slept.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und die Mädchen schlafen, und ich neue Mädchen schlafen.", "score": 0.0}685{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_489.wav", "doc_id": "SUkmfOTvGi.seg_489", "src_text": "And this shows us that adaptive overfitting in this case is not observed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Und das zeigt uns, dass in diesem Fall keine anpassungsfähige Überdimensionierung beobachtet wird.", "score": 70.0}686{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_450.wav", "doc_id": "hgIDlKNiFM.seg_450", "src_text": "In total, we have seven models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Insgesamt haben wir sieben Modelle.", "score": 100.0}687{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_109.wav", "doc_id": "uZBWfYjYnf.seg_109", "src_text": "If we go on and we receive another speech chunk, and our model predicts other three words and we will look at those cross-attention weights, we will see that no word points to the last lambda speech frames.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wenn wir fortfahren und einen anderen Sprechrhythmus erhalten, und unser Modell weitere drei Wörter vorhersagt, werden wir auf diese Cross-Attention-Ways schauen. Wir werden sehen, dass keine Worte auf die letzten Lamda-Sprechrahmen verweisen.", "score": 90.0}688{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_140.wav", "doc_id": "wLqFAuDnKa.seg_140", "src_text": "We saw that the actual form of the prompting doesn't have a big influence in the case of several short promptings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir sahen, dass die tatsächliche Form der Anrufung keinen großen Einfluss auf den Fall von mehreren Anrufungen hat.", "score": 60.0}689{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_558.wav", "doc_id": "rISrKoXQCx.seg_558", "src_text": "So some preliminary results demonstrate that first, language models do have varying political leanings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "vorläufige Ergebnisse, dass die ersten Sprachmodelle immer noch unterschiedliche politische Präferenzen aufweisen.", "score": 75.0}690{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_569.wav", "doc_id": "rISrKoXQCx.seg_569", "src_text": "So this indicates that language models can also pick up the polarisation in our society.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "so dass die Sprachmodelle auch die Polarisierung in unserer Gesellschaft aufgreifen können.", "score": 85.0}691{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_450.wav", "doc_id": "hgIDlKNiFM.seg_450", "src_text": "In total, we have seven models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Insgesamt haben wir sieben Modelle.", "score": 100.0}692{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_95.wav", "doc_id": "uZBWfYjYnf.seg_95", "src_text": "And what are the problems of the current SimulST models?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Und was sind die Probleme der aktuellen SimulST-Modelle?", "score": 100.0}693{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_851.wav", "doc_id": "GvEBWkLmuI.seg_851", "src_text": "And more broadly, dominant groups in society are both linguistically and socially unmarked, while the marginalized groups are usually marked.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "markiert. Und mehr oder weniger, die dominierenden Gruppen in der Gesellschaft sind sowohl sprachlich als auch sozial unmarkiert, während die marginalisierten Gruppen üblicherweise markiert sind. Unsere Methode", "score": 94.0}694{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_627.wav", "doc_id": "oeooqChmKK.seg_627", "src_text": "To summarize the main takeaways of our paper, many coreference resolution models appear unable to reason over knowledge from different sources without task-specific training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Um die wichtigsten Aspekte unseres Papiers zusammenzufassen: Viele Korrelationsmodelle scheinen nicht in der Lage zu sein, Wissen aus verschiedenen Quellen ohne spezifisches Training zu nutzen.", "score": 64.0}695{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_855.wav", "doc_id": "GvEBWkLmuI.seg_855", "src_text": "So first we use a lexicon of stereotypes, and we find that the generated personas contain a lot more stereotypes than the human-written ones.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "also verwenden wir zunächst ein Elektronen-Stereotyp und stellen fest, dass die geborene Person eine viel größere Anzahl an Stereotypen enthält als die der Menschen, die sie kennen.", "score": 75.0}696{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_608.wav", "doc_id": "oeooqChmKK.seg_608", "src_text": "And second, background knowledge such as \"Judges decide cases in law courts.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diener ein Richter ist, und zweitens, die Hintergrundkenntnis, wie etwa, dass Richter Fälle in Gerichtshöfen entscheiden.", "score": 43.0}697{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_603.wav", "doc_id": "oeooqChmKK.seg_603", "src_text": "Servin and Kea met at a park.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Servin und Kiah trafen sich nach einem", "score": 50.0}698{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_876.wav", "doc_id": "GvEBWkLmuI.seg_876", "src_text": "And finally, there should really be increased transparency about bias mitigation methods, because for instance, like these positive stereotypes, we don't know if it's because there is some sort of weird overly-excessive value alignment going on, or maybe some other anti-stereotyping methods that are resulting in these pernicious patterns.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und schließlich sollten die Transparenz über die Bias-Methode wirklich erhöht werden. Weil wir zum Beispiel nicht wissen, ob es auf diese positiven Stereotypen etwas wie „irgendwie seltsam“ gibt. Übermäßig hohe Werte werden aufgenommen, oder vielleicht andere, wie anti-stereotypierende Methoden, die zu diesen schädlichen Mustern führen.", "score": 97.0}699{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_74.wav", "doc_id": "TVCREhgqUP.seg_74", "src_text": "Conceptually, our permutation model works roughly like this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Konzeptionell funktioniert unser Permutationsmodell ungefähr so.", "score": 90.0}700{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_274.wav", "doc_id": "PIZEXUFLAR.seg_274", "src_text": "For investigating multi-modal instruction tuning on our proposed dataset, we take OFA, a unified multi-modal pre-trained model, as our base model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "der multimodalen Anweisung in unserem Vorschlag verwenden wir Ofa als ein einheitliches multimodales Darstellungsmodell als Basismodell.", "score": 100.0}701{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_877.wav", "doc_id": "GvEBWkLmuI.seg_877", "src_text": "We just really can't make any assumptions or really study that further, without more transparency.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das ist wirklich nicht möglich, oder man muss das weiter mit mehr Transparenz untersuchen.", "score": 50.0}702{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_78.wav", "doc_id": "TVCREhgqUP.seg_78", "src_text": "We determine the third token in the output in a similar way by jumping to another multiset token.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir bestimmen das dritte Token in der Ausgabe auf ähnliche Weise, indem wir zu einem anderen Multisets-Token springen.", "score": 90.0}703{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_679.wav", "doc_id": "oaOHnMCwad.seg_679", "src_text": "Where prospective API is able to detect correctly toxic instances.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "denn Ihre API ist in der Lage, korrekte toxische Einflüsse zu erkennen.", "score": 50.0}704{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_194.wav", "doc_id": "SLpqvupgvW.seg_194", "src_text": "For example, the same genre or the same artist for a song.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "beispielsweise dasselbe Genre oder dasselbe Künstler. Wenn wir diese alternative", "score": 70.0}705{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_817.wav", "doc_id": "WTTtiRKFZI.seg_817", "src_text": "Here we have coordination of two verbs and there's no outsides, external governor.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "es eine Koordination von zwei Wörtern gibt, aber keine Außen-Gouverneur.", "score": 100.0}706{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_740.wav", "doc_id": "XejEJmgUmE.seg_740", "src_text": "We're trying to revisit the MPP pipeline by asking the model to evaluate acceptability on longer and longer sequences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir versuchen, die NPB-Pipeline zu überprüfen, indem wir das Modell bitten, die Akzeptabilität auf längeren und längeren Sequenzen zu bewerten.", "score": 75.0}707{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_821.wav", "doc_id": "WTTtiRKFZI.seg_821", "src_text": "So I'll concentrate on the right one.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "uns auf die rechte Spalte zu konzentrieren.", "score": 4.0}708{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_772.wav", "doc_id": "WTTtiRKFZI.seg_772", "src_text": "As you may know, there are different dependency structures assumed by different theories and corpus approaches.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wie wir wissen, gibt es verschiedene Abhängigkeitsstrukturen, die von verschiedenen Theorien und Körperprozessen genutzt werden,", "score": 75.0}709{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_161.wav", "doc_id": "SLpqvupgvW.seg_161", "src_text": "My name is Javad Hosseini and this is a joint work with Filip Radlinski, Silvia Pareti, and Annie Louis.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Mein Name ist Jawad Hussain, und dies ist eine gemeinsame Arbeit mit Philip Radlinski, Sylvia Parati und Annie Tows.", "score": 60.0}710{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_728.wav", "doc_id": "XejEJmgUmE.seg_728", "src_text": "Hi, everyone.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo alle,", "score": 100.0}711{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_173.wav", "doc_id": "SLpqvupgvW.seg_173", "src_text": "We're not aware of a larger-scale public data set for the task, so we collect one using crowd annotation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "einen öffentlichen Datensatz a large-scale public dataset für die Aufgabe gibt, also sammeln wir einen mithilfe der Crowdannotation. Unsere", "score": 71.0}712{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_203.wav", "doc_id": "SLpqvupgvW.seg_203", "src_text": "Here are some examples from our dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hier sind einige Beispiele aus unserem Datensatz.", "score": 100.0}713{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_141.wav", "doc_id": "wLqFAuDnKa.seg_141", "src_text": "It's crucial for zero and one-shot prompting.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Es ist entscheidend für Null und einen", "score": 50.0}714{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_423.wav", "doc_id": "WBLMIsdIrq.seg_423", "src_text": "This again demonstrates that it is difficult to determine the best document-level translation system if we use corpus-level metrics alone.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Dies zeigt erneut, dass es schwierig ist, das beste Dokumenten-Niveau-Übersetzungs-System zu bestimmen, wenn wir nur Korpus-Niveau-Metriken verwenden.", "score": 95.0}715{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_567.wav", "doc_id": "rISrKoXQCx.seg_567", "src_text": "We separately pretrain language models on the two different temporal corpora.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "nach dem Vierzigsten Präsidenten und dem Fünfzigsten Präsidenten der Vereinigten Staaten.", "score": 0.0}716{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_213.wav", "doc_id": "SLpqvupgvW.seg_213", "src_text": "Here is a link to our dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist ein Link zu unserem Datensatz,", "score": 90.0}717{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_158.wav", "doc_id": "wLqFAuDnKa.seg_158", "src_text": "Thank you very much.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Vielen Dank.", "score": 100.0}718{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_822.wav", "doc_id": "WTTtiRKFZI.seg_822", "src_text": "What we see here is that when the governor is on the left, the tendency for the left conjunct to be shorter grows steadily, with the absolute difference in words, and the same is observed when there is no governor as in coordination of sentences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Was wir sagen, ist, dass das so ist, wenn der Regierungschef auf der linken Seite ist. Die Tendenz, dass der linke Konjunktur kürzer ist, wächst stetig mit dem absoluten Unterschied in den Worten, und das gleiche wird beobachtet, wenn es keinen Gouverneur gibt, der die Sätze koordiniert,", "score": 60.0}719{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_601.wav", "doc_id": "oeooqChmKK.seg_601", "src_text": "Servin is a judge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Serwin ist ein", "score": 20.0}720{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_30.wav", "doc_id": "aQpIWggfCo.seg_30", "src_text": "Since large language models are costly to deploy, it's essential to enable language planning ability of smaller and specialized models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Da große Sprachmodelle teuer zu deployen sind, ist es wichtig, Sprachplanung mit etwas kleineren und spezialisierten Modellen zu ermöglichen.", "score": 90.0}721{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_41.wav", "doc_id": "aQpIWggfCo.seg_41", "src_text": "In summary, we establish the constrained language planning problem.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zusammenfassend: Wir stellen das Problem der konstrizierten Sprachplanung fest;", "score": 90.0}722{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_598.wav", "doc_id": "oeooqChmKK.seg_598", "src_text": "We introduce a coreference resolution task, designed to probe for the ability to draw on knowledge available in different sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir führen eine Korreferenzlösungsaufgabe ein, die darauf ausgelegt ist, die Fähigkeit zu testen, auf Wissen aus verschiedenen Quellen zu ziehen.", "score": 90.0}723{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_697.wav", "doc_id": "oaOHnMCwad.seg_697", "src_text": "We then take the annotations by demographic and compare them to the models and datasets using a Pearson's R correlation score, and thus our framework actually differs from annotator disagreement literature by comparing end users with models and datasets, predictions and labels, as opposed to looking at just annotator agreement or modelling annotator distributions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir nehmen dann die demographischen Annotationen und vergleichen sie mit den Modellen und Datensätzen, indem wir die Korrelationskennzahlen verwenden. Daher unterscheidet sich unser Framework von der Annotatoren-Disagreement-Literatur, indem wir Endnutzer mit Modellen und Datensätzen, Vorhersagen und Etiketten vergleichen, anstatt nur eine Annotatoren-Übereinstimmung oder Modellierung von Annotatoren-Verteilungen zu betrachten.", "score": 86.0}724{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_465.wav", "doc_id": "SUkmfOTvGi.seg_465", "src_text": "Let's get started.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "2023? Lassen Sie", "score": 0.0}725{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_805.wav", "doc_id": "WTTtiRKFZI.seg_805", "src_text": "Right?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Ordnung, aber", "score": 0.0}726{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_356.wav", "doc_id": "gGbuDbHhyc.seg_356", "src_text": "Second, if clean data is required, or if clean data is mandatory for WSL to work, then how many clean samples do we need?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "saubere Daten erforderlich sind oder wenn saubere Daten für die Arbeit von WSL obligatorisch sind, wie viele saubere Stichproben benötigen wir dann?", "score": 75.0}727{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_271.wav", "doc_id": "PIZEXUFLAR.seg_271", "src_text": "Therefore, this motivates us to build a multi-modal instruction tuning dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Instruction-Set, daher motiviert es uns, ein Multimodal Instruction-Set zu erstellen.", "score": 75.0}728{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_540.wav", "doc_id": "dvGkKzmIaN.seg_540", "src_text": "We also validate the covertness of the provided embedding by visualising the embedding of sentences on four dataset [INAUDIBLE 4:39] PCA.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir haben auch die Vertraulichkeit der bereitgestellten Embeddings durch die Visualisierung der Embeddings von Sätzen gefordert, z.B.", "score": 66.0}729{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_332.wav", "doc_id": "dJGfOSFgZO.seg_332", "src_text": "These reliable, informative, and distinct ABC-Eval metrics enable us to evaluate conversational AI with a higher resolution than previous methods are able to achieve.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese zuverlässigen, aussagekräftigen und klaren A.B.C.-Maße ermöglichen es uns, die Konversation mit einer höheren Auflösung zu bewerten, als es mit den vorherigen Methoden möglich war.", "score": 85.0}730{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_579.wav", "doc_id": "rISrKoXQCx.seg_579", "src_text": "So a little bit of discussion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "die Verwendung von politischen Modellierungen der Sprache entstehen, angehen sollten.", "score": 25.0}731{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_604.wav", "doc_id": "oeooqChmKK.seg_604", "src_text": "After a long day at work deciding cases in a law court, he was happy to relax.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "langen Arbeitstag im Park, um sich zu entspannen.", "score": 69.0}732{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_63.wav", "doc_id": "TVCREhgqUP.seg_63", "src_text": "This can be complicated and sometimes a computationally expensive process.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dies kann kompliziert sein und manchmal ein computergestütztes Verfahren", "score": 96.0}733{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_484.wav", "doc_id": "SUkmfOTvGi.seg_484", "src_text": "To our next question, what causes the performance drop of some models, We had two hypothesis.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zur nächsten Frage: Was verursacht den Leistungseinbruch einiger Modelle? Wir hatten zwei Hypothesen:", "score": 100.0}734{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_323.wav", "doc_id": "dJGfOSFgZO.seg_323", "src_text": "To determine what kind of evaluation is most effective, we selected four state-of-the-art chat models and evaluated them on 100 human-bot conversations per model using ABC-Eval.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Um herauszufinden, was für eine Art von Bewertung die effektivste ist, haben wir vier Chat-Modelle ausgewählt und sie anhand von Hunderten von menschlichen Gesprächen pro Modell bewertet.", "score": 85.0}735{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_478.wav", "doc_id": "SUkmfOTvGi.seg_478", "src_text": "The first one is the model architecture.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die erste ist die Modellarchitektur.", "score": 100.0}736{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_661.wav", "doc_id": "FLkGnzVRew.seg_661", "src_text": "Next, to improve the number of dissonance examples, we use a Probability-of-Rare-Class strategy — PRC — to select mostly the examples that are highly likely to be descended by the current model at any round of rare.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "nächstes verbessern wir die Anzahl der Dissimilitud-Beispiele, indem wir die Wahrscheinlichkeit einer seltenen Klasse-Strategie PRC verwenden, um die Beispiele zu wählen, die am wahrscheinlichsten von dem aktuellen Modell in jeder Runde von ALE abweichen.", "score": 100.0}737{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_727.wav", "doc_id": "oaOHnMCwad.seg_727", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 100.0}738{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_553.wav", "doc_id": "rISrKoXQCx.seg_553", "src_text": "On the other hand, these different political opinions are inherently socially biased and might lead to potential fairness issues in downstream task applications.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und auf der anderen Seite sind diese unterschiedlichen politischen Meinungen sozial voreingenommen und möglicherweise zu potenziellen Fairnessproblemen in downstream-Anwendungen.", "score": 90.0}739{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_701.wav", "doc_id": "oaOHnMCwad.seg_701", "src_text": "We host 2 tasks on lab in the wild, one of them being social acceptability, and the way this works is that participants will read a situation from the social chemistry dataset and, then they'll write how socially acceptable a situation is.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Leben im Freien, einer davon ist die soziale Akzeptanz, und die Art und Weise, wie dies funktioniert, ist, dass die Teilnehmer eine Situation aus dem Social Chemistry-Datensatz lesen und dann beurteilen, wie sozial akzeptabel eine Situation ist.", "score": 75.0}740{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_152.wav", "doc_id": "wLqFAuDnKa.seg_152", "src_text": "The insights that we gained from the human evaluation that we performed using the MQM framework said that the fluency of PaLM is comparable to state-of-the-art systems but the main difference comes from the accuracy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Erkenntnisse, die wir aus der menschlichen Bewertung gewinnen, die wir mit dem MKM-Framework durchführen, sind, dass die Flüssigkeit von Palms mit dem Stand der Kunstsysteme vergleichbar ist, aber der Hauptunterschied kommt von der Genauigkeit, insbesondere.", "score": 67.0}741{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_368.wav", "doc_id": "gGbuDbHhyc.seg_368", "src_text": "Finally, the performance improvement claimed in previous WSL approaches can be easily achieved by allowing to continue fine-tuning on the clean validation samples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Schließlich kann die Leistungsverbesserung, die bei früheren WSL-Ansätzen behauptet wurde, leicht erzielt werden, indem man das Feintuning an sauberen Validierungsbeispielen fortsetzt.", "score": 100.0}742{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_857.wav", "doc_id": "GvEBWkLmuI.seg_857", "src_text": "So, while the generated personas have much higher rates of the lexicon words, the human-written ones have a much wider distribution of words, while the stereotype words that are in the generated personas are really just the words \"tall\" and \"athletic\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "haben also viel höhere Raten der Luxon-Wörter, die humanen haben eine viel breitere Verteilung der Wörter, während die stereotypen Wörter in den generierten Personen wirklich nur die Wörter sind. Es sind", "score": 24.0}743{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_11.wav", "doc_id": "aQpIWggfCo.seg_11", "src_text": "Since no dataset of specific goals exists to support our study, we have to acquire these goals first.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "gibt keine Daten außerhalb von spezifischen Zielen, um unsere Studienzeit zu verlängern. Wir müssen diese Ziele zuerst erwerben.", "score": 67.0}744{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_184.wav", "doc_id": "SLpqvupgvW.seg_184", "src_text": "The second one, which is the alternative question is generated as follows.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die zweite, die alternative Frage, wird wie folgt generiert.", "score": 100.0}745{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_858.wav", "doc_id": "GvEBWkLmuI.seg_858", "src_text": "So, really just only the positive or at least non-negative ones.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Also wirklich nur die positiven oder zumindest keine negativen.", "score": 90.0}746{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_734.wav", "doc_id": "XejEJmgUmE.seg_734", "src_text": "Which can also include grammaticality like BLiMP, SyntaxGym, or acceptability in terms of stereotypes such as CrowS pairs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "auch Grammatikalität wie „blimp“ oder Akzeptanz in Bezug auf Stereotypen wie „Cousins“ beinhalten können.", "score": 67.0}747{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_596.wav", "doc_id": "oeooqChmKK.seg_596", "src_text": "Therefore, successful models for knowledge-intensive NLU tasks require the ability to integrate and use both pretrain-time and inference-time knowledge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "erfordern erfolgreiche Modelle für kognitiv intensive NLU-Aufgaben die Fähigkeit, Vorwissenszeit und Inferenzzeit zu integrieren und zu nutzen.", "score": 100.0}748{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_159.wav", "doc_id": "SLpqvupgvW.seg_159", "src_text": "Hi!", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo,", "score": 67.0}749{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_815.wav", "doc_id": "WTTtiRKFZI.seg_815", "src_text": "So the governor is on the left in this example \"I saw Bart and Lisa\" so is the governor is on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "auf der linken Seite, und der Gouverneur ist auf der linken Seite.", "score": 30.0}750{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_349.wav", "doc_id": "gGbuDbHhyc.seg_349", "src_text": "In weakly supervised learning, training algorithms are proposed to robustly train neural networks under such label noise so that the trained models still generalize well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "generalisieren. Bei der schwachen Beobachtungstraining werden Trainingsalgorithmen vorgeschlagen, um Nervennetze unter solchen Etiketten „Noise“ robust zu trainieren, so dass die Trainingsmodelle immer noch stark generalisiert werden.", "score": 49.0}751{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_601.wav", "doc_id": "oeooqChmKK.seg_601", "src_text": "Servin is a judge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Serwin ist Richter,", "score": 100.0}752{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_727.wav", "doc_id": "oaOHnMCwad.seg_727", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 99.0}753{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_384.wav", "doc_id": "WBLMIsdIrq.seg_384", "src_text": "A Data-driven, Multilingual Exploration\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "durch Sprachverarbeitung bezieht.", "score": 0.0}754{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_376.wav", "doc_id": "gGbuDbHhyc.seg_376", "src_text": "For example, report if the model selection is done via clean validation samples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ob die Modellauswahl mit sauberen Validierungsmustern durchgeführt wird.", "score": 90.0}755{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_449.wav", "doc_id": "hgIDlKNiFM.seg_449", "src_text": "Another also based on CamemBERT, but trained this time on the 4 GB of clinical notes and finally, one based on English biomedical model PubMedBERT, and trained on 4 GB of set of NACHOS.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ebenfalls auf Camembert, aber trainierte diesmal auf vier Kilogramm von Kinkanlot. Und schließlich haben wir ein Modell auf der Grundlage eines englischen biomedizinischen Modells, Bumet, und trainieren es auf vier Gigabyte von Naturstoffen, insgesamt haben", "score": 75.0}756{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_596.wav", "doc_id": "oeooqChmKK.seg_596", "src_text": "Therefore, successful models for knowledge-intensive NLU tasks require the ability to integrate and use both pretrain-time and inference-time knowledge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "erfordern erfolgreiche Modelle für wissensintensive NLU-Aufgaben die Fähigkeit, sowohl Vortrainingszeit als auch Inferenzzeit zu integrieren und zu verwenden.", "score": 100.0}757{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_57.wav", "doc_id": "TVCREhgqUP.seg_57", "src_text": "In this example, the model has seen shallow recursion during training and is tested on an example with deeper recursion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In diesem Beispiel hat das Modell eine flache Rückwärtsbewegung während des Trainings und wird auf einem Beispiel mit tiefer Rückwärtsbewegung getestet.", "score": 99.0}758{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_209.wav", "doc_id": "SLpqvupgvW.seg_209", "src_text": "If the language model has access to some partially overlapping background knowledge, then the accuracy is between 82 to 87%, which is more realistic.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wenn das Sprachmodell auf einige teilweise überlappende Hintergrundkenntnisse zugreifen kann, liegt die Genauigkeit zwischen achtundachtzig und neunundsiebzig Prozent, was beispielsweise dann", "score": 55.0}759{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_174.wav", "doc_id": "SLpqvupgvW.seg_174", "src_text": "Our data set covers three different domains: music, books, and recipes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Datensätze umfassen drei verschiedene Domänen: Musik, Bücher und Rezepte.", "score": 100.0}760{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_367.wav", "doc_id": "gGbuDbHhyc.seg_367", "src_text": "As we can see, if we have 10 samples per class, direct fine-tuning starts to beat WSL approaches.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wie wir sehen können, wenn wir zehn Proben pro Klasse haben, beginnt die Direkt-Fine-Tuning zu WSL-Ansätzen.", "score": 100.0}761{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_51.wav", "doc_id": "TVCREhgqUP.seg_51", "src_text": "In the context of semantic parsing, testing for compositional generalization might look like this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Im Kontext der semantischen Analyse könnte das Testen für die Kompositionsgeneralisierung wie folgt aussehen:", "score": 86.0}762{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_344.wav", "doc_id": "gGbuDbHhyc.seg_344", "src_text": "I'd like to begin with a brief introduction to weak supervision and weakly supervised learning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ich würde gerne mit einer kurzen Einführung zum wöchentlichen Überwachung und wöchentlichen Überwachungsunterricht beginnen.", "score": 19.0}763{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_62.wav", "doc_id": "TVCREhgqUP.seg_62", "src_text": "This works well, but trees are usually not given and need to be obtained somehow.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das funktioniert auch, aber es ist normalerweise nicht möglich, etwas davon zu erhalten.", "score": 96.0}764{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_397.wav", "doc_id": "WBLMIsdIrq.seg_397", "src_text": "To answer the first question, we started by measuring how much a word depends on context during translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Um die erste Frage zu beantworten, begannen wir damit, zu messen, wie viel ein Wort bei der Übersetzung von Kontext abhängt.", "score": 98.0}765{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_441.wav", "doc_id": "hgIDlKNiFM.seg_441", "src_text": "However, French didn't have any open source model for biomedical until now.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "neues Open-Source-Modell für BioMedicine. Wir stellen uns also die Frage,", "score": 72.0}766{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_338.wav", "doc_id": "dJGfOSFgZO.seg_338", "src_text": "We hope ABC-Eval can be leveraged by others in the field as a meaningful step in this direction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir hoffen, dass A.B.C.Eval von anderen in diesem Bereich als bedeutender Schritt in diese Richtung angesehen wird, und", "score": 100.0}767{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_463.wav", "doc_id": "SUkmfOTvGi.seg_463", "src_text": "Hello everyone, my name is Shuheng.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hallo alle, ich heiße Suhun.", "score": 100.0}768{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_839.wav", "doc_id": "GvEBWkLmuI.seg_839", "src_text": "Immediately we see that, while the outputs aren't overtly negative or toxic in the traditional sense of these words, there are some interesting patterns.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Sofort sehen wir, dass die Ausgaben nicht offensichtlich negativ oder giftig sind, im traditionellen Sinne dieser Wörter. Es gibt einige interessante Muster:", "score": 69.0}769{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_766.wav", "doc_id": "XejEJmgUmE.seg_766", "src_text": "That is, when we perturb the sentences in the acceptable domain, we see similar increase in all the perturbations and when we perturb the sentences in the unacceptable domain, we see decrease in MPP judgments in similar fashion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn wir die Sätze im akzeptablen Bereich stören, sehen wir einen ähnlichen Anstieg aller Störungen, und wenn wir die Sätze im unakzeptablen Bereich stören, sehen wir einen ähnlichen Rückgang der MP-P-Judikate. So", "score": 100.0}770{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_319.wav", "doc_id": "dJGfOSFgZO.seg_319", "src_text": "We call this approach annotating behaviors in chat or ABC-Eval in short.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir nennen diese Herangehensweise Annotieren von Verhaltensweisen im Chat oder ABC für Kurzform.", "score": 100.0}771{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_658.wav", "doc_id": "FLkGnzVRew.seg_658", "src_text": "Next, we determine the best method to update a model with new data from each round of active learning and annotations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Als nächstes bestimmen wir die beste Methode zur Aktualisierung eines Modells mit neuen Daten aus jeder Runde der aktiven Lern- und Annotierungen: kumulative Akkumulationen aller", "score": 66.0}772{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_499.wav", "doc_id": "SUkmfOTvGi.seg_499", "src_text": "Thank you so much.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 100.0}773{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_522.wav", "doc_id": "dvGkKzmIaN.seg_522", "src_text": "Before these main steps, we first select a trigger set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hauptschritte ausführen, wählen wir zunächst einen Auslöser. Der", "score": 71.0}774{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_51.wav", "doc_id": "TVCREhgqUP.seg_51", "src_text": "In the context of semantic parsing, testing for compositional generalization might look like this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Im Kontext des semantischen Parsings, bei dem man für die generelle Zusammensetzung testet, sieht", "score": 100.0}775{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_686.wav", "doc_id": "oaOHnMCwad.seg_686", "src_text": "And as a researcher, positionality can influence the research process and its outcomes and results because it can change the decisions that researchers make.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "als Forscher kann die Positionalität den Forschungsprozess und seine Ergebnisse und Ergebnisse beeinflussen, weil sie die Entscheidungen, die Forscher treffen, verändern kann.", "score": 76.0}776{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_159.wav", "doc_id": "SLpqvupgvW.seg_159", "src_text": "Hi!", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo,", "score": 100.0}777{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_587.wav", "doc_id": "rISrKoXQCx.seg_587", "src_text": "I think that's pretty much all I have for today.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ich glaube, das ist sehr viel, ich bin gestorben, ich bin fünfmal für heute gestorben,", "score": 80.0}778{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_559.wav", "doc_id": "rISrKoXQCx.seg_559", "src_text": "They occupy all four quadrants on the political campus.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Man kann", "score": 0.0}779{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_627.wav", "doc_id": "oeooqChmKK.seg_627", "src_text": "To summarize the main takeaways of our paper, many coreference resolution models appear unable to reason over knowledge from different sources without task-specific training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Um die Hauptaspekte unseres Papiers zusammenzufassen: Viele Referenzmodelle scheinen nicht in der Lage zu sein, Wissen aus verschiedenen Quellen ohne taskspezifische Schulung zu verarbeiten.", "score": 98.0}780{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_71.wav", "doc_id": "TVCREhgqUP.seg_71", "src_text": "That's why in the second step we use another model to predict a permutation to put them into the right order.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Deshalb verwenden wir in der zweiten Phase ein anderes Modell, um eine Permutation vorherzusagen, um sie in die richtige Reihenfolge zu bringen.", "score": 100.0}781{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_661.wav", "doc_id": "FLkGnzVRew.seg_661", "src_text": "Next, to improve the number of dissonance examples, we use a Probability-of-Rare-Class strategy — PRC — to select mostly the examples that are highly likely to be descended by the current model at any round of rare.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Um die Anzahl der Beispiele zu erhöhen, wählen wir die Beispiele aus, die am ehesten durch das aktuelle Modell unterschieden werden können.", "score": 5.0}782{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_790.wav", "doc_id": "WTTtiRKFZI.seg_790", "src_text": "Right?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "weil hier", "score": 0.0}783{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_123.wav", "doc_id": "wLqFAuDnKa.seg_123", "src_text": "This is joint work with my colleagues from Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dies ist eine gemeinsame Arbeit mit meinen Kolleginnen und Kollegen von Google Translate.", "score": 100.0}784{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_418.wav", "doc_id": "WBLMIsdIrq.seg_418", "src_text": "We then use the MuDA tagger, by applying the tagger on a parallel corpus that we want to use for evaluation and we apply our translation metrics of choice on the context-dependent examples that the MuDA tagger has identified.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dann verwenden wir den Mudatagger, indem wir den Tagger auf einem parallelen Korpus anwenden, den wir für die Bewertung verwenden möchten, und wir wenden unsere Übersetzungsmaße der Wahl auf die kontextabhängigen Beispiele an, die der Mudatagger identifiziert hat.", "score": 100.0}785{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_686.wav", "doc_id": "oaOHnMCwad.seg_686", "src_text": "And as a researcher, positionality can influence the research process and its outcomes and results because it can change the decisions that researchers make.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und als Forscher kann die Positionierung den Forschungsprozess und seine Ergebnisse beeinflussen, weil sie die Entscheidungen der Forscher ändern kann.", "score": 99.0}786{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_356.wav", "doc_id": "gGbuDbHhyc.seg_356", "src_text": "Second, if clean data is required, or if clean data is mandatory for WSL to work, then how many clean samples do we need?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "saubere Daten erforderlich sind oder wenn saubere Daten für die Arbeit von WSL obligatorisch sind, wie viele saubere Beispiele benötigen wir dann?", "score": 100.0}787{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_158.wav", "doc_id": "wLqFAuDnKa.seg_158", "src_text": "Thank you very much.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "vielen Dank.", "score": 100.0}788{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_137.wav", "doc_id": "wLqFAuDnKa.seg_137", "src_text": "So, it's important to select a good prompting strategy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "erreichen, also ist es wichtig, eine gute Promotionsstrategie auszuwählen.", "score": 75.0}789{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_685.wav", "doc_id": "oaOHnMCwad.seg_685", "src_text": "This is a concept widely used in critical studies, specifically in feminist and queer academic spaces.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Dies ist ein weit verbreitetes Konzept in kritischen Studien, insbesondere in feministischen und queer akademischen Räumen,", "score": 100.0}790{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_234.wav", "doc_id": "oYCKgTzTDy.seg_234", "src_text": "We also test Monolingual Few-shot setting by training monolingual models with only 10% of training data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Außerdem testen wir die Einstellungen für die Monolinguale, indem wir Monolinguale Modelle mit nur dreizehn Prozent der Trainingsdaten trainieren.", "score": 85.0}791{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_377.wav", "doc_id": "gGbuDbHhyc.seg_377", "src_text": "Second, WSL approaches should be compared with few-shot learning baselines, as both work on clean samples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zweitens sollten WSAL-Ansätze mit zukünftigen Lernbasen verglichen werden, eine vorgesehene Arbeit an klaren Mustern;", "score": 85.0}792{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_766.wav", "doc_id": "XejEJmgUmE.seg_766", "src_text": "That is, when we perturb the sentences in the acceptable domain, we see similar increase in all the perturbations and when we perturb the sentences in the unacceptable domain, we see decrease in MPP judgments in similar fashion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "sind. Das heißt, wenn wir die Sätze in der akzeptablen Domäne stören, sehen wir einen ähnlichen Anstieg aller Störungen und wenn wir die Sätze in der nicht akzeptablen Domäne stören, sehen wir einen ähnlichen Rückgang der MP-Richtlinien. Der", "score": 83.0}793{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_156.wav", "doc_id": "wLqFAuDnKa.seg_156", "src_text": "And that's it for this really short overview.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und das ist es, was wir für diese wirklich schockierende Überprüfung", "score": 65.0}794{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_477.wav", "doc_id": "SUkmfOTvGi.seg_477", "src_text": "Throughout experiments we found that there are three main ingredients that are needed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "unsere Experimente haben wir festgestellt, dass es drei Hauptbestandteile gibt, die benötigt werden.", "score": 99.0}795{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_381.wav", "doc_id": "gGbuDbHhyc.seg_381", "src_text": "Please feel free to check it out.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "finden. Bitte fühlen Sie sich frei, ihn", "score": 73.0}796{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_315.wav", "doc_id": "dJGfOSFgZO.seg_315", "src_text": "Therefore, you might want to evaluate multiple dimensions of chat quality to understand the strengths and weaknesses of the model on a finer-grained level.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "daher möchte ich, dass Sie mehrere Dimensionen der Dialogqualität bewerten, um die Stärken und Schwächen des Modells zu verstehen.", "score": 85.0}797{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_400.wav", "doc_id": "WBLMIsdIrq.seg_400", "src_text": "In this work, we extend CXMI to Pointwise CXMI which can measure context usage at the sentence level or at the word level.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In dieser Arbeit erweitern wir CxMI zu point-wise CxMI, mit dem man Kontextnutzung am Satzlevel messen kann. Level, oder auf Wortebene:", "score": 100.0}798{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_743.wav", "doc_id": "XejEJmgUmE.seg_743", "src_text": "So for example, here we have chosen like a typical pair of grammaticality from the BLiMP data set from the Adjunct Island case.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daher haben wir beispielsweise hier wie eine typische Paarung von Grammatikalität aus dem Datenbestand von der Adjuvant Island-Kase ausgewählt.", "score": 81.0}799{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_373.wav", "doc_id": "gGbuDbHhyc.seg_373", "src_text": "Their performance gain and practicality are heavily overestimated.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ihr Leistungsgewinn und ihre Praktikabilität werden stark überschätzt.", "score": 100.0}800{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_855.wav", "doc_id": "GvEBWkLmuI.seg_855", "src_text": "So first we use a lexicon of stereotypes, and we find that the generated personas contain a lot more stereotypes than the human-written ones.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die wir verwenden, und wir stellen fest, dass die Stereotypen, die wir verwenden, viel mehr Stereotypen enthalten als die Stereotypen, die", "score": 81.0}801{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_766.wav", "doc_id": "XejEJmgUmE.seg_766", "src_text": "That is, when we perturb the sentences in the acceptable domain, we see similar increase in all the perturbations and when we perturb the sentences in the unacceptable domain, we see decrease in MPP judgments in similar fashion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "reagieren. Das bedeutet, wenn wir die Sätze in der akzeptablen Domäne stören, sehen wir eine ähnliche Zunahme der Störungen, und wenn wir die Sätze in der unakzeptablen Domäne stören, sehen wir eine ähnliche Abnahme der Urteile.", "score": 99.0}802{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_481.wav", "doc_id": "SUkmfOTvGi.seg_481", "src_text": "We found that usually larger models lead to better generalization.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir haben festgestellt, dass größere Modelle zu einer besseren Generalisierung führen.", "score": 100.0}803{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_640.wav", "doc_id": "FLkGnzVRew.seg_640", "src_text": "So why does this matter?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Warum ist das wichtig?", "score": 95.0}804{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_694.wav", "doc_id": "oaOHnMCwad.seg_694", "src_text": "The first step is to re annotate data sets with diverse annotators.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Der erste Schritt besteht darin, Datensätze mit verschiedenen Annotatoren neu zu annotieren.", "score": 99.0}805{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_850.wav", "doc_id": "GvEBWkLmuI.seg_850", "src_text": "So when people are describing a warrior who is a woman, they'll usually actually specify \"woman warrior\" and mark the term with \"woman\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "einen Krieger beschreibt, der normalerweise eine Frau ist, dann ist das normalerweise eine Frau. Und mehr noch, die dominierenden", "score": 25.0}806{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_106.wav", "doc_id": "uZBWfYjYnf.seg_106", "src_text": "A word is emitted if the attention is not concentrated, that is, its sum is below a certain threshold alpha towards the last lambda speech frames, meaning that the received information is enough stable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Ein Wort wird ausgesprochen, wenn die Spannung nicht konzentriert ist, d. h. wenn ihr Wert unter einem bestimmten Schwellenwert Alpha liegt, was bedeutet, dass die erhaltenen Informationen nicht stabil sind.", "score": 90.0}807{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_646.wav", "doc_id": "FLkGnzVRew.seg_646", "src_text": "We used dissonance-first approach, as seen in the flow chart here.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir verwendeten den ersten Ansatz zur Diskrepanz, wie er hier im Flussdiagramm dargestellt", "score": 80.0}808{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_791.wav", "doc_id": "WTTtiRKFZI.seg_791", "src_text": "Because here between the verb and the direct object is an adjunct: \"yesterday\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "zwischen dem Verb und dem direkten Objekt gestern Abend noch etwas hinzugekommen ist.", "score": 75.0}809{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_834.wav", "doc_id": "GvEBWkLmuI.seg_834", "src_text": "To overcome these limitations, we rely on the property that these newer instruction-tuned LLMs are very good at responding to instructions and prompts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Diese Einschränkungen zu überschreiten, sind wir auf die Eigenschaft angewiesen, dass diese neuen Anweisungen sehr gut auf Anweisungen antworten. So kann", "score": 75.0}810{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_624.wav", "doc_id": "oeooqChmKK.seg_624", "src_text": "When trained on KITMUS, however, both C2F and BERT4Coref perform significantly better than the random choice.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir jedoch auf Kidmus ausgebildet werden, funktionieren sowohl Sea to Earth als auch Bert Forquerth deutlich besser als die Durandal-Option.", "score": 80.0}811{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_263.wav", "doc_id": "PIZEXUFLAR.seg_263", "src_text": "Hello everyone, my name is Ying and my colleague Zhiyang and I will be presenting our research on MultiInstruct improving Multi-Modal Zero-Shot Learning via Instruction Tuning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo alle, mein Name ist Ying und mein Kollege Jian und ich werden unsere Forschung über Multi-Instruct, Verbesserung des multimodalen sozialen Lernens durch Instruktion, vorstellen.", "score": 90.0}812{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_625.wav", "doc_id": "oeooqChmKK.seg_625", "src_text": "This suggests that when trained on generic reference resolution data sets, most learn to exploit surface cues, which are not useful when testing on KITMUS where such queues have been removed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies lässt darauf schließen, dass Mäuse, wenn sie auf allgemeine Quervergleichsdatensätze trainiert werden, Oberflächenmerkmale ausnutzen, die bei der Überprüfung in einem Käfig nicht nützlich sind.", "score": 80.0}813{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_152.wav", "doc_id": "wLqFAuDnKa.seg_152", "src_text": "The insights that we gained from the human evaluation that we performed using the MQM framework said that the fluency of PaLM is comparable to state-of-the-art systems but the main difference comes from the accuracy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die Erkenntnisse, die wir aus der menschlichen Analyse gewinnen, und die wir mit dem MQR-Verfahren durchführen, ist, dass die Fließfähigkeit von Palmen mit dem Zustand der Kunstsysteme vergleichbar ist, aber der Hauptunterschied kommt aus der Genauigkeit. Insbesondere,", "score": 90.0}814{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_497.wav", "doc_id": "SUkmfOTvGi.seg_497", "src_text": "We hope our paper calls for more research on how to improve generalizations of the models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir hoffen, dass unsere Arbeit mehr Forschung auf den Weg bringt, wie man die Modelle verbessern kann.", "score": 90.0}815{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_450.wav", "doc_id": "hgIDlKNiFM.seg_450", "src_text": "In total, we have seven models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Insgesamt haben wir sieben Modelle.", "score": 100.0}816{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_266.wav", "doc_id": "PIZEXUFLAR.seg_266", "src_text": "However, most previous works on instruction tuning focused on improving the zero-shot performance on language only tasks, while computer vision and multi-modal tasks have been left out.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die meisten früheren Arbeiten zum Anpassen von Anweisungen konzentrierten sich jedoch auf die Verbesserung der sekundären Leistung bei Sprachaufgaben, wobei Computerseh- und Multimodaltasks ausgelassen wurden.", "score": 45.0}817{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_107.wav", "doc_id": "uZBWfYjYnf.seg_107", "src_text": "For example, if we receive a speech chunk containing \"I'm going to talk about...\" and our model predicts the translation in German, and we will look at the cross-attention weights, we'll see that the first two words points to the earliest received speech frames, while the last word points to the last received speech frames, as lambda speech frames.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zum Beispiel, wenn wir einen Textschnipsel erhalten, der „Ich werde darüber sprechen“ enthält, und unser Modell die Übersetzung ins Deutsche vorhersagt. Und wir werden auf die Kreuzbelastung achten. Wir werden sehen, dass die ersten beiden Wörter auf die frühesten empfangenen Sprachrahmen verweisen, während das letzte Wort auf die zuletzt empfangenen Sprachrahmen, die sogenannten „Lamda“-Sprachrahmen, verweist.", "score": 75.0}818{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_349.wav", "doc_id": "gGbuDbHhyc.seg_349", "src_text": "In weakly supervised learning, training algorithms are proposed to robustly train neural networks under such label noise so that the trained models still generalize well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "in WSL Weekly Superwise Learning ist", "score": 5.0}819{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_414.wav", "doc_id": "WBLMIsdIrq.seg_414", "src_text": "So now we use our findings from our analysis to design a benchmark for document-level translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "verwenden wir unsere Ergebnisse aus unseren Analysen, um einen Benchmark für die Dokumenten-Novelle zu entwerfen.", "score": 40.0}820{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_685.wav", "doc_id": "oaOHnMCwad.seg_685", "src_text": "This is a concept widely used in critical studies, specifically in feminist and queer academic spaces.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "häufig verwendet wird, insbesondere in feministischen und queer akademischen Räumen.", "score": 3.0}821{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_84.wav", "doc_id": "TVCREhgqUP.seg_84", "src_text": "First of all, the alignment between input and output is not given in the training data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "vor. Zunächst einmal ist die Ausrichtung zwischen Input und Output in den Trainingsdaten nicht angegeben.", "score": 95.0}822{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_617.wav", "doc_id": "oeooqChmKK.seg_617", "src_text": "Here's an example of how we control the availability of facts in the true sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hier ist ein Beispiel dafür, wie wir die Verfügbarkeit von Fakten aus wahren Quellen kontrollieren.", "score": 95.0}823{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_113.wav", "doc_id": "uZBWfYjYnf.seg_113", "src_text": "But also we want that they are shifted on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "diesem Plot ist. Aber wir wollen auch, dass sie nach links verschoben", "score": 60.0}824{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_791.wav", "doc_id": "WTTtiRKFZI.seg_791", "src_text": "Because here between the verb and the direct object is an adjunct: \"yesterday\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "zwischen dem Verb und dem Objekt ein Abstand ist. Und", "score": 96.0}825{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_780.wav", "doc_id": "WTTtiRKFZI.seg_780", "src_text": "The conjunction headed approach assumed in Prague dependency treebanks, where coordinate structures are headed by the conjunction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Prag-Ansatz, dem Konjunktur-Ansatz, der in Prag-Abhängigkeitstrinzenen durchgeführt wird, wobei Koordinatursysteme von der Konjunktur durchgeführt werden.", "score": 95.0}826{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_70.wav", "doc_id": "TVCREhgqUP.seg_70", "src_text": "After the first step, we have all the right tokens, but they're not ordered.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Nach dem ersten Schritt haben wir alle richtigen Tokens, aber sie sind nicht bestellt.", "score": 90.0}827{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_483.wav", "doc_id": "SUkmfOTvGi.seg_483", "src_text": "Here we also found that more fine tuning examples, actually also leads to better generalization.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hier haben wir auch festgestellt, dass mehr Feintuning-Beispiele tatsächlich zu einer besseren Generalisierung führen.", "score": 100.0}828{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_442.wav", "doc_id": "hgIDlKNiFM.seg_442", "src_text": "So we ask ourselves a question about what is the most appropriate data sources for a wide range of usage and those crawled data are good substitution for clinical data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "was die geeignetsten Datenquellen für eine Vielzahl von Anwendungen sind, und diese Daten sind gute Ersatzdaten für klinische Daten.", "score": 90.0}829{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_698.wav", "doc_id": "oaOHnMCwad.seg_698", "src_text": "Our frame is largely enabled through Lab in the Wild and online crowdsourcing platform for where HCI collaborator.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "weitgehend über Lab in the Wild, eine Online-Crowdsourcing-Plattform, nutzbar. In", "score": 30.0}830{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_467.wav", "doc_id": "SUkmfOTvGi.seg_467", "src_text": "We observe that models have been used in CoNLL-2003 to develop NER for almost 20 years and this naturally raises several problems.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir stellen fest, dass Modelle, die Carnell 2003 zur Entwicklung von NER verwendet haben, fast 20 Jahre lang verwendet wurden, und dies wirft natürlich viele Probleme auf:", "score": 75.0}831{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_162.wav", "doc_id": "SLpqvupgvW.seg_162", "src_text": "Our goal is to understand users’ language when they want to make a choice.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Unser Ziel ist es, die Sprache der Benutzer zu verstehen, wenn sie eine Wahl treffen möchten.", "score": 100.0}832{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_27.wav", "doc_id": "aQpIWggfCo.seg_27", "src_text": "We only keep the script if the target goal scores the highest in the goal set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir behalten nur das Skript bei, wenn das Zielziel der höchsten Bewertung im Zielziel erreicht.", "score": 78.0}833{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_479.wav", "doc_id": "SUkmfOTvGi.seg_479", "src_text": "Through our experiments we found that the transformer models normally generalize better to new data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Durch unsere Experimente haben wir festgestellt, dass sich die Transformatormodelle normalerweise besser an neue Daten anpassen.", "score": 100.0}834{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_569.wav", "doc_id": "rISrKoXQCx.seg_569", "src_text": "So this indicates that language models can also pick up the polarisation in our society.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "bewegt, nach 2017, was darauf hindeutet, dass Sprachmodelle auch die Polarisierung in unserer Gesellschaft aufgreifen können.", "score": 99.0}835{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_450.wav", "doc_id": "hgIDlKNiFM.seg_450", "src_text": "In total, we have seven models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir sieben Modelle.", "score": 31.0}836{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_811.wav", "doc_id": "WTTtiRKFZI.seg_811", "src_text": "So when the difference between the lengths of the two conjuncts grows, the shorter conjunct prefers to be the first one, stronger, right?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Längendifferenzs wächst, also wenn die Länge des Längendifferenzs wächst, bevorzugt der kürzere Konjunkt zuerst der stärkere ist, also", "score": 10.0}837{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_444.wav", "doc_id": "hgIDlKNiFM.seg_444", "src_text": "Afterwards, we ask ourselves how much data do we need to train a specialized model on French data?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wie viel Daten benötigen wir, um ein spezialisiertes Modell auf französischen Daten zu trainieren?", "score": 95.0}838{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_681.wav", "doc_id": "oaOHnMCwad.seg_681", "src_text": "Where prospective AP is really not as sensitive to offensive terms that are more common in Indian contexts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Perspektive der AP wirklich nicht auf offensichtliche Begriffe in indischen Kontexten gerichtet ist.", "score": 62.0}839{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_17.wav", "doc_id": "aQpIWggfCo.seg_17", "src_text": "Results in the figure show that the semantic completeness in generated scripts is acceptable but the faithfulness to the constraints cannot be guaranteed.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "in der Grafik zeigen, dass die semantische Vollständigkeit in generierten Skripten akzeptabel ist, aber die Treue zu den Einschränkungen kann nicht garantiert werden.", "score": 97.0}840{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_233.wav", "doc_id": "oYCKgTzTDy.seg_233", "src_text": "In this setting, the source language is the same as target language, for example German to German or English to English.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In dieser Einstellung ist die Quellensprache dieselbe wie die Zielsprache, zum Beispiel Deutsch zu Deutsch oder Englisch zu Englisch.", "score": 100.0}841{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_854.wav", "doc_id": "GvEBWkLmuI.seg_854", "src_text": "Now for some results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Für Ergebnisse, also", "score": 67.0}842{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_627.wav", "doc_id": "oeooqChmKK.seg_627", "src_text": "To summarize the main takeaways of our paper, many coreference resolution models appear unable to reason over knowledge from different sources without task-specific training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Um die wichtigsten Schlussfolgerungen des Berichts zusammenzufassen, Viele Koherenzmodellierungsmodelle scheinen ohne task-spezifische Ausbildung nicht in der Lage zu sein, Überwissens aus verschiedenen Quellen zu begründen.", "score": 75.0}843{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_206.wav", "doc_id": "SLpqvupgvW.seg_206", "src_text": "Results with T5 XL model are summarized below.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Ergebnisse mit dem großen Modell T5-XL werden unten zusammengefasst.", "score": 66.0}844{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_737.wav", "doc_id": "XejEJmgUmE.seg_737", "src_text": "The current MPP pipeline basically doesn't allow us to evaluate a model's acceptance towards longer sentences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "für MPP-Modelle erlaubt uns eigentlich nicht, die Akzeptanz eines Modells gegenüber längeren Sätzen zu bewerten.", "score": 85.0}845{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_879.wav", "doc_id": "GvEBWkLmuI.seg_879", "src_text": "Have a good time at ACL.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zuhören, ich hatte eine gute Zeit.", "score": 24.0}846{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_544.wav", "doc_id": "dvGkKzmIaN.seg_544", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir kommen,", "score": 0.0}847{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_648.wav", "doc_id": "FLkGnzVRew.seg_648", "src_text": "As can be seen here, dissonance was only found in 3.5% of the annotated pairs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wie man hier sehen kann, war die Diskrepanz nur in fünf Prozent der annotierten Paare zu", "score": 85.0}848{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_125.wav", "doc_id": "wLqFAuDnKa.seg_125", "src_text": "It's trained on a large collection of text, comprising 780 billion tokens.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Es basiert auf einer großen Textkollektion, die 780 Milliarden Dokumente", "score": 98.0}849{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_485.wav", "doc_id": "SUkmfOTvGi.seg_485", "src_text": "The first one is adaptive overfitting, which is overfitting costs by reusing the same test set over and over again and this is usually manifested as the diminishing returns on a new test set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "die erste ist die adaptive Überanpassung, die durch die wiederholte Verwendung desselben Tests verursacht wird, und dies zeigt sich normalerweise, wenn die Abnahme auf dem neuen Test zurückkehrt.", "score": 75.0}850{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_141.wav", "doc_id": "wLqFAuDnKa.seg_141", "src_text": "It's crucial for zero and one-shot prompting.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist entscheidend für Null- und eine Anregung, aber", "score": 9.0}851{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_204.wav", "doc_id": "SLpqvupgvW.seg_204", "src_text": "For example, \"the one without words\", \"not the one with the 12 year old boy\", or \"the fictional one\", or \"comes from Azerbaijan\", and so on.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "zum Beispiel der ohne Worte, nicht der mit dem zwölfjährigen Jungen, oder der fiktionale oder aus Aserbaidschan. Der Alternativenkorpus hat", "score": 68.0}852{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_272.wav", "doc_id": "PIZEXUFLAR.seg_272", "src_text": "Here we present MultiInstruct, the first multi-modal instruction tuning benchmark dataset that consists of 62 diverse multi-modal tasks covering 10 broad categories.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hier stellen wir MultiInstruct vor, das erste multimodale Benchmark-Datensatz, der aus 62 verschiedenen multimodalen Aufgaben besteht, die 10 Kategorien abdecken.", "score": 97.0}853{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_590.wav", "doc_id": "oeooqChmKK.seg_590", "src_text": "This work is a collaboration between McGill University, Mila, and Microsoft Research.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Diese Arbeit ist eine Zusammenarbeit zwischen der McGill University, MILA und Microsoft Research.", "score": 100.0}854{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_153.wav", "doc_id": "wLqFAuDnKa.seg_153", "src_text": "So, in particular, the most common errors are omission errors.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "häufig sind Auslassungsfehler. Es scheint, dass", "score": 40.0}855{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_259.wav", "doc_id": "oYCKgTzTDy.seg_259", "src_text": "And our results show many interesting findings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "und unsere Ergebnisse zeigen viele interessante Ergebnisse,", "score": 100.0}856{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_121.wav", "doc_id": "uZBWfYjYnf.seg_121", "src_text": "Thanks for your attention.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "vielen Dank für Ihre Aufmerksamkeit.", "score": 100.0}857{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_462.wav", "doc_id": "hgIDlKNiFM.seg_462", "src_text": "So thank you for this presentation, and we are looking forward to exchange at the poster session in Toronto.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dank für diese Präsentation, und wir freuen uns darauf, in der Post Office in Toronto zu handeln.", "score": 22.0}858{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_706.wav", "doc_id": "oaOHnMCwad.seg_706", "src_text": "Our study in the end amassed over 16,000 annotations from over 1000 annotators from 87 countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und am Ende sammelten über 16.000 Annotierungen von über 1.000 Annotatoren aus 87", "score": 99.0}859{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_233.wav", "doc_id": "oYCKgTzTDy.seg_233", "src_text": "In this setting, the source language is the same as target language, for example German to German or English to English.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In diesem Sinne ist die Quelle die gleiche wie die Ziel-Sprache, zum Beispiel Deutsch zu Deutsch oder Englisch zu Englisch.", "score": 99.0}860{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_820.wav", "doc_id": "WTTtiRKFZI.seg_820", "src_text": "So we showed that by measuring length in characters, the first column, in syllables the middle column, and in words the right column.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir, dass wir, indem wir die Länge von Zeichen messen, dem ersten Wort in Sätzen, dem mittleren Wort in Sätzen und den Worten im Text, dem", "score": 25.0}861{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_6.wav", "doc_id": "aQpIWggfCo.seg_6", "src_text": "Planning for the goals with specific constraints, such as \"make a chocolate cake\", still remains under-studied.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Planung für Ziele mit spezifischen Zielen, die spezifischen Einschränkungen wie das Backen eines Schokoladenkuchens unterliegen, ist immer noch nicht ausreichend erforscht.", "score": 75.0}862{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_524.wav", "doc_id": "dvGkKzmIaN.seg_524", "src_text": "We assume the provider can collect a general text corpus and count the word frequency with it.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir gehen davon aus, dass der Anbieter einen allgemeinen Textkorpus sammeln und die Wortfrequenz zählen kann. Bei der", "score": 98.0}863{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_571.wav", "doc_id": "rISrKoXQCx.seg_571", "src_text": "So we see that if we investigate the per category performance, that is to say if we separate the performance into different demographics or political leaning of news media we can see a pattern.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wir die Leistungskategorie untersuchen, das bedeutet, wenn wir die Leistung aufteilen. Verschiedene Demografien oder politische Nachrichtenmedien können ein Muster erkennen,", "score": 92.0}864{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_727.wav", "doc_id": "oaOHnMCwad.seg_727", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 100.0}865{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_792.wav", "doc_id": "WTTtiRKFZI.seg_792", "src_text": "However, this effect may be ameliorated when the direct object is very heavy and very long.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "das Objekt des Direkts ist ein sehr schweres und sehr langes Objekt, weil es dann", "score": 79.0}866{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_76.wav", "doc_id": "TVCREhgqUP.seg_76", "src_text": "For the first output position, we simply select one, as highlighted in red.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Für die erste Ausgangsposition wählen wir einfach eines, das rot markiert ist.", "score": 95.0}867{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_711.wav", "doc_id": "oaOHnMCwad.seg_711", "src_text": "We find that Dynahate is also most aligned to English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "auch heraus, dass die Datenmodelle für die meisten englischsprachigen Länder am besten geeignet sind.", "score": 7.0}868{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_535.wav", "doc_id": "dvGkKzmIaN.seg_535", "src_text": "We compute the similarity difference between benign and backdoor data set which is defined as delta cosine and delta L2.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir berechnen den Ähnlichkeitsunterschied zwischen dem normalen und dem Hintertür-Datensatz, der als Delta-Kosinus und Delta-L-Zwei definiert ist.", "score": 70.0}869{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_346.wav", "doc_id": "gGbuDbHhyc.seg_346", "src_text": "Instead, we label the data using weak labeling sources, such as simple heuristic rules, knowledge bases, or low-quality crowdsourcing, as illustrated in the figure on the right.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "schwache Beschriftungsquellen, wie einfache heuristische Regeln, Wissensbasen oder niedrigwertige Cloud-Quellen, wie es in der Abbildung rechts dargestellt ist.", "score": 70.0}870{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_439.wav", "doc_id": "hgIDlKNiFM.seg_439", "src_text": "Since then, this model has been adapted to many other languages, like in French with CamemBERT, and also in domains like biomedical with PubMedBERT and BioBERT and on clinical with ClinicalBERT, but mostly in English.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Seitdem wurde dieses Modell auf viele andere Sprachen wie Französisch mit Camembert und andere Domänen wie Biomedizin mit Pametber und Biober übernommen, und auf klinisch mit klinisch übernommen, aber meistens auf Englisch.", "score": 50.0}871{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_696.wav", "doc_id": "oaOHnMCwad.seg_696", "src_text": "And so we opt to re annotate data to get many annotates for instance and to get a rich set of demographic data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir entscheiden uns also dafür, die Daten zu reannotieren, um viele Anwender beispielsweise zu erhalten und einen reichen Satz an demographischen Daten zu erhalten.", "score": 91.0}872{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_98.wav", "doc_id": "uZBWfYjYnf.seg_98", "src_text": "And training and maintaining several models to reach different latency regimes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und man trainiert und hält mehrere Modelle, um verschiedene Latenzregime zu ermitteln,", "score": 85.0}873{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_143.wav", "doc_id": "wLqFAuDnKa.seg_143", "src_text": "It's the examples that carry most of the weight.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "sind die Beispiele, die den größten", "score": 66.0}874{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_837.wav", "doc_id": "GvEBWkLmuI.seg_837", "src_text": "And we can immediately see that this is very generalizable to any demographic because we can just specify whatever identity marker that we want into this prompt.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und wir können sofort sehen, dass das sehr allgemein für jede Demografie ist, weil wir einfach jede Identität angeben können, die wir in diesem Prom haben wollen.", "score": 83.0}875{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_146.wav", "doc_id": "wLqFAuDnKa.seg_146", "src_text": "In particular, we compare the selecting prompts from the training data for the WMT evaluations on the dev data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "insbesondere, dass wir die Auswahlanregungen aus den Trainingsdaten der WMT-Evaluierungen oder den Testdaten vergleichen.", "score": 81.0}876{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_621.wav", "doc_id": "oeooqChmKK.seg_621", "src_text": "We evaluate the data set both with human study participants, and established coreference resolution models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir bewerten sowohl die Datensätze mit den menschlichen Studienteilnehmern als auch die etablierten Lösungsmodelle.", "score": 93.0}877{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_437.wav", "doc_id": "hgIDlKNiFM.seg_437", "src_text": "And finally, we conclude about the experiments and give you more details about how to access those models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und schließen schließlich über die Experimente ab und geben Ihnen mehr Details darüber, wie Sie die Modelle zugreifen können.", "score": 71.0}878{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_596.wav", "doc_id": "oeooqChmKK.seg_596", "src_text": "Therefore, successful models for knowledge-intensive NLU tasks require the ability to integrate and use both pretrain-time and inference-time knowledge.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "hat. Daher erfordern erfolgreiche Modelle für wissensintensive LU-Aufgaben die Fähigkeit, sowohl vorbereitete Zeit als auch Inferenzzeit zu nutzen.", "score": 70.0}879{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_266.wav", "doc_id": "PIZEXUFLAR.seg_266", "src_text": "However, most previous works on instruction tuning focused on improving the zero-shot performance on language only tasks, while computer vision and multi-modal tasks have been left out.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Allerdings konzentrierten sich die meisten vorherigen Arbeiten zur Anweisungstuning auf die Verbesserung der Null-Schussleistung bei Sprachaufgaben, wobei Computer-Vision- und multimodale Aufgaben außer Acht gelassen wurden.", "score": 97.0}880{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_170.wav", "doc_id": "SLpqvupgvW.seg_170", "src_text": "Or when the user wants to specify a preference.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Oder wenn der Benutzer eine Präferenz angeben möchte,", "score": 100.0}881{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_831.wav", "doc_id": "GvEBWkLmuI.seg_831", "src_text": "However, these measures have various limitations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese Maßnahmen haben jedoch verschiedene Einschränkungen, sie", "score": 100.0}882{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_600.wav", "doc_id": "oeooqChmKK.seg_600", "src_text": "Here is an example from our data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "ein Beispiel aus unserem Datensatz:", "score": 100.0}883{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_282.wav", "doc_id": "PIZEXUFLAR.seg_282", "src_text": "We use all the instances in the test split for each task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "verwenden alle Instanzen im Testsplit für", "score": 100.0}884{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_45.wav", "doc_id": "aQpIWggfCo.seg_45", "src_text": "Thanks for your time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Vielen Dank für Ihre Zeit.", "score": 100.0}885{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_223.wav", "doc_id": "oYCKgTzTDy.seg_223", "src_text": "The Lambda calculus is missing, or they're only evaluated on certain neural models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Der Mond ist sichtbar. Oder sie werden nur anhand bestimmter neuerer Modelle bewertet.", "score": 14.0}886{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_196.wav", "doc_id": "SLpqvupgvW.seg_196", "src_text": "So what we do is that we show some background knowledge about the two entities.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir zeigen also einige Hintergrundwissen über die beiden Entitäten.", "score": 98.0}887{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_296.wav", "doc_id": "PIZEXUFLAR.seg_296", "src_text": "Here we can see, as the amount of task increases, the model achieves better performance and in the meantime, lower sensitivity.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hier können wir sehen, dass, wenn die Anzahl der Aufgaben zunimmt, das Modell eine bessere Leistung erreicht und gleichzeitig eine geringere Empfindlichkeit.", "score": 100.0}888{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_306.wav", "doc_id": "PIZEXUFLAR.seg_306", "src_text": "So this is a QR code for our data and model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist dies ein QR-Code für unsere Daten und unser Modell.", "score": 100.0}889{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_808.wav", "doc_id": "WTTtiRKFZI.seg_808", "src_text": "So what we did, we extracted various statistics about coordination from the enhanced version of the Penn Treebank and see the paper \"Why wouldn't you use universal dependencies\" and these statistics confirm the observation made many times before that left conjuncts tend to be shorter.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir haben verschiedene Statistiken über die Koordination aus der erweiterten Version der Pentribank und sehen uns das Papier an, warum wir keine universellen Abhängigkeiten verwenden würden. Und diese Statistiken bestätigen die Beobachtung, dass die linken Konjunktionen tendenziell kürzer sind", "score": 80.0}890{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_515.wav", "doc_id": "dvGkKzmIaN.seg_515", "src_text": "Finally, the watermark needs to be transferable to the attacker's services during the model extraction process.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Schließlich muss das Wasserzeichen während des Modellentnahmeverfahrens auf die Oberfläche des Angreifers übertragen werden.", "score": 85.0}891{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_177.wav", "doc_id": "SLpqvupgvW.seg_177", "src_text": "In the first bubble, Bob says, \"Remember that song we were listening to yesterday?\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In der ersten Blase sagt Bob „Denk an das Lied, das wir gestern Abend gehört haben“.", "score": 90.0}892{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_695.wav", "doc_id": "oaOHnMCwad.seg_695", "src_text": "And we ought to do this over looking at the demographics of original data sets annotators, because, usually only a few annotators annotate each instance and because demographics are rarely collected and shared.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und wir werden das in den Demografien der Originaldatensätze nachschlagen, Annotatoren, weil normalerweise nur wenige Annotatoren vorhanden sind und weil die Demografien tatsächlich gesammelt und geteilt werden.", "score": 50.0}893{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_393.wav", "doc_id": "WBLMIsdIrq.seg_393", "src_text": "And some people have suggested targeted evaluation on context-dependent translations, but these resources only support limited types of context-dependent translations and limited sets of languages since they usually rely on domain knowledge and human curation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "einige Leute haben vorgeschlagen, kontextabhängige Übersetzungen auf kontextabhängige Übersetzungen zu bewerten, aber diese Ressourcen unterstützen nur begrenzte Arten von kontextabhängigen Übersetzungen und begrenzte Sprachmengen, da sie normalerweise auf Domänwissen und menschliche Kuration angewiesen sind.", "score": 86.0}894{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_412.wav", "doc_id": "WBLMIsdIrq.seg_412", "src_text": "And finally, we look at different individual tokens that have high P-CXMI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und schließlich sehen wir uns unterschiedliche individuelle Token an, die eine hohe PSXMI haben,", "score": 61.0}895{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_80.wav", "doc_id": "TVCREhgqUP.seg_80", "src_text": "To give you a teaser of the experimental results, here we compare our method with other treeless models on the COGS benchmark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Um Ihnen einen Überblick über die experimentellen Ergebnisse zu geben, vergleichen wir hier unsere Methode mit anderen Baumlos-Modellen auf der Grundlage des", "score": 75.0}896{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_367.wav", "doc_id": "gGbuDbHhyc.seg_367", "src_text": "As we can see, if we have 10 samples per class, direct fine-tuning starts to beat WSL approaches.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wie wir sehen können, beginnt Direct Fine-Tuning, wenn wir zehn Proben pro Klasse haben, WS-Ansätze zu schlagen.", "score": 78.0}897{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_792.wav", "doc_id": "WTTtiRKFZI.seg_792", "src_text": "However, this effect may be ameliorated when the direct object is very heavy and very long.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dieser Effekt kann jedoch verbessert werden, wenn das Zielobjekt sehr schwer und sehr lang ist,", "score": 93.0}898{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_118.wav", "doc_id": "uZBWfYjYnf.seg_118", "src_text": "And we also see that if we consider the actual elapsed time or the computational-aware time, that is the fastest strategy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und wir sehen auch, dass, wenn wir die tatsächliche Laufzeit oder die computergestützte Arbeitszeit betrachten, die FASTER-Strategie die schnellste ist.", "score": 57.0}899{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_388.wav", "doc_id": "WBLMIsdIrq.seg_388", "src_text": "Well, if the previous sentence was \"Things could start to get dangerous if the ministers find out\", then \"mole\" refers to a spy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn die vorherige Aussage lautete, dass die Dinge gefährlich werden könnten, wenn die Minister es herausfinden, bezieht sich „Moe“ auf einen Spion.", "score": 95.0}900{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_510.wav", "doc_id": "dvGkKzmIaN.seg_510", "src_text": "To protect the copyright of embedding as services, one of the solutions is to embed a watermark in the provider service and detect whether another service contain the watermark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Um die Urheberrechte von Embedding-Services zu schützen, kann eine der Lösungen darin bestehen, ein Wasserzeichen in den Diensten des Anbieters zu embedden und zu überprüfen, ob ein anderer Dienst das Wasserzeichen enthält.", "score": 67.0}901{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_456.wav", "doc_id": "hgIDlKNiFM.seg_456", "src_text": "Overall, from-scratch pre-training seems to obtain higher performance on most of the tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Insgesamt scheint das Scratch-Training eine höhere Leistung bei den meisten Aufgaben zu erzielen. Allerdings können unsere Experimente,", "score": 85.0}902{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_212.wav", "doc_id": "SLpqvupgvW.seg_212", "src_text": "We've also shown that the models are domain-generalizable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir haben auch gezeigt, dass die Modelle domänenspezifisch sind, hier", "score": 1.0}903{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_75.wav", "doc_id": "TVCREhgqUP.seg_75", "src_text": "We go from left to right over the output and determine which multiset token to put in every position.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir gehen von links nach rechts über die Ausgabe und bestimmen, welchen Multisets-Token wir in jede Position setzen.", "score": 85.0}904{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_59.wav", "doc_id": "TVCREhgqUP.seg_59", "src_text": "In particular, they often fail to reproduce the systematic correspondences between input and output, such as those that are color-coded in the example.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Insbesondere funktionieren die systematischen Korrespondenzen zwischen Input und Output nicht immer. Die", "score": 80.0}905{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_275.wav", "doc_id": "PIZEXUFLAR.seg_275", "src_text": "OFA uses a unified vocabulary for language, image tokens and the coordinates of a bounding box.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "OFA verwendet eine einheitliche Vokabularität für Sprache, Bildsymbolen und den Koordinator einer Bindungskiste.", "score": 47.0}906{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_292.wav", "doc_id": "PIZEXUFLAR.seg_292", "src_text": "So this measures the model's ability to consistently produce the same outputs for the same task regardless of the slight variation in the wording of the instruction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "die Fähigkeit des Modells zu messen, die gleichen Ergebnisse für die gleiche Aufgabe zu produzieren, unabhängig von geringfügigen Abweichungen in der Formulierung der Anweisung.", "score": 80.0}907{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_774.wav", "doc_id": "WTTtiRKFZI.seg_774", "src_text": "So in this case, Lisa.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "in diesem Fall Ilsa. Ähnliche Ansätze", "score": 74.0}908{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_197.wav", "doc_id": "SLpqvupgvW.seg_197", "src_text": "For songs, we simply show a Google search link to each song and then ask the annotators to listen to at least some of each song, and read about each song.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "zeigen. Für Songs zeigen wir einfach einen Google-Suchlink zu jedem Song. Und bitten Sie dann die Annotatoren, zumindest einige der Lieder anzuhören und darüber zu lesen.", "score": 100.0}909{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_676.wav", "doc_id": "oaOHnMCwad.seg_676", "src_text": "This work was done in collaboration with some folks at the University of Washington and the Allen Institute for AI, namely Sebastian Santy, Ronan Le Bras, Katharina Reinecke and Maarten Sap.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Arbeit wurde in Zusammenarbeit mit einigen Kolleginnen und Kollegen der Universität Washington und des AI-Instituts der Universität Washington durchgeführt, darunter Sebastian Santy, Ronan Labras, Caterina Rinaea und Martin Sap.", "score": 90.0}910{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_540.wav", "doc_id": "dvGkKzmIaN.seg_540", "src_text": "We also validate the covertness of the provided embedding by visualising the embedding of sentences on four dataset [INAUDIBLE 4:39] PCA.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir validierten auch die Verschlüsselung der bereitgestellten Einbettung, indem wir die Einbettung von Sätzen auf Virtuelle-Z-V-P-A validierten.", "score": 50.0}911{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_421.wav", "doc_id": "WBLMIsdIrq.seg_421", "src_text": "But then if we use COMET, context-aware models perform best.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Aber wenn wir kommentierte, kontextsensitive Modelle", "score": 70.0}912{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_15.wav", "doc_id": "aQpIWggfCo.seg_15", "src_text": "We find that all language models achieve unsatisfactory results on planning for specific goals.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir stellen fest, dass alle Linearmodelle bei der Planung für bestimmte Ziele unzufriedenstellende Ergebnisse liefern.", "score": 80.0}913{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_432.wav", "doc_id": "hgIDlKNiFM.seg_432", "src_text": "In this presentation, we first talk about language modeling in healthcare.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In diesem Vortrag sprechen wir zunächst über Sprachmodellierung im Gesundheitswesen, dann", "score": 100.0}914{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_333.wav", "doc_id": "dJGfOSFgZO.seg_333", "src_text": "You can see that in the results of our experiment that several challenges still remain and have been precisely quantified.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "den Ergebnissen unseres Experiments können Sie sehen, dass mehrere Herausforderungen noch bestehen und präzise quantifiziert wurden.", "score": 95.0}915{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_543.wav", "doc_id": "dvGkKzmIaN.seg_543", "src_text": "That's all.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das ist alles, danke.", "score": 100.0}916{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_458.wav", "doc_id": "hgIDlKNiFM.seg_458", "src_text": "Which is not the case for the model based on CamemBERT weights and tokenizer, which suffer from stability issues.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies ist nicht der Fall für das Modell, das auf Kammanberwägen basiert und Stabilitätsprobleme aufweist.", "score": 60.0}917{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_113.wav", "doc_id": "uZBWfYjYnf.seg_113", "src_text": "But also we want that they are shifted on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "diesem Plot sind. Aber auch wir wollen, dass sie auf der linken Seite stehen.", "score": 95.0}918{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_348.wav", "doc_id": "gGbuDbHhyc.seg_348", "src_text": "If we directly train neural networks on weakly labeled data, the neural networks tend to memorize the label noise and do not generalize.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "mit menschlichen Notationen vergleicht, sind die schwachen Notationen immer noch ein allgemeiner Trend. In jüngsten Arbeiten", "score": 50.0}919{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_42.wav", "doc_id": "aQpIWggfCo.seg_42", "src_text": "We evaluate constrained language planning ability of large language models and develop an over-generate-then-filter method for large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "bewerten die eingeschränkte Sprachplanungsfähigkeit großer Sprachmodelle und entwickeln eine übergenerierende Filtermethode für große Sprachmodelle.", "score": 95.0}920{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_152.wav", "doc_id": "wLqFAuDnKa.seg_152", "src_text": "The insights that we gained from the human evaluation that we performed using the MQM framework said that the fluency of PaLM is comparable to state-of-the-art systems but the main difference comes from the accuracy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die Erkenntnisse, die wir aus der menschlichen Analyse gewonnen haben, die wir mit dem MQM-Rahmenwerk durchgeführt haben, sind, dass die Flüssigkeit von Palm mit dem Zustand der Systeme vergleichbar ist, aber die Hauptunterschiede kommen von der Genauigkeit.", "score": 90.0}921{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_570.wav", "doc_id": "rISrKoXQCx.seg_570", "src_text": "So last but not least, we evaluate language models with different political leanings on hate speech detection and fake news detection to NLP applications that often involve language models and could have very significant implications.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir bewerten Sprachmodelle mit unterschiedlichen politischen Ausrichtungen, Sprachprüfungen und Nachrichtenprüfungen, die Sprachmodelle beinhalten können und sehr signifikante Implikationen haben. Also sagen wir das, wenn wir", "score": 70.0}922{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_586.wav", "doc_id": "rISrKoXQCx.seg_586", "src_text": "Ok, great.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Okay, großartig,", "score": 100.0}923{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_513.wav", "doc_id": "dvGkKzmIaN.seg_513", "src_text": "Second, the watermark should not degrade the utility of the provided embeddings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zweitens sollte das Watermark den Nutzen der bereitgestellten Embeddings nicht beeinträchtigen.", "score": 100.0}924{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_185.wav", "doc_id": "SLpqvupgvW.seg_185", "src_text": "We always use a simple template.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir verwenden immer eine einfache Vorlage:", "score": 95.0}925{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_835.wav", "doc_id": "GvEBWkLmuI.seg_835", "src_text": "So we can ask the model to generate a persona, which is a depiction of an imagined individual using a prompt like \"Imagine you are an Asian woman.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "das Modell eine Person erzeugen, die eine asiatische Frau beschreibt, die so aussieht, als ob sie sich selbst beschreiben würde.", "score": 75.0}926{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_64.wav", "doc_id": "TVCREhgqUP.seg_64", "src_text": "Typically, this involves considerable formalism-specific pre-processing of the logical forms, for example, to handle variable symbols.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Typischerweise beinhaltet dies eine erhebliche formalismusspezifische Präprozessierung der logischen Formen, zum Beispiel zur Verarbeitung von variablen Symbolen. Auch spezielle Grammatikanalysen", "score": 90.0}927{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_573.wav", "doc_id": "rISrKoXQCx.seg_573", "src_text": "And vice versa, right-leaning language models are better at detecting hate speech targeting white and men, however worse at detecting hate speech targeting at black LGBTQ plus and other minority communities.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "treffen. Und umgekehrt: Modellierende Sprachmodelle sind besser darin, Heusprech zu erkennen, die auf Weiß und Mann zielen, aber es ist besser, Heusprech zu erkennen, die auf Schwarz, LGBQT+ und andere Minderheitsgemeinschaften zielen.", "score": 40.0}928{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_103.wav", "doc_id": "uZBWfYjYnf.seg_103", "src_text": "And leverage the knowledge already acquired by the model through the attention mechanism between audio input and textual output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und die Kenntnisse, die durch das Modell über den Spannungsmechanismus zwischen Audioeingabe und Text Ausgabe bereits erworben wurden,", "score": 60.0}929{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_528.wav", "doc_id": "dvGkKzmIaN.seg_528", "src_text": "The weight of the target embedding is proportional to the number of triggers in the sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Gewicht des Ziel-Embeddings ist proportional zur Anzahl der Auslöser in einem Satz.", "score": 89.0}930{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_83.wav", "doc_id": "TVCREhgqUP.seg_83", "src_text": "In our paper, we solve a couple of interesting technical challenges.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In unserem Papier stellen wir einige interessante technische Herausforderungen", "score": 60.0}931{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_600.wav", "doc_id": "oeooqChmKK.seg_600", "src_text": "Here is an example from our data set.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hier ist ein Beispiel aus unserem Datensatz:", "score": 100.0}932{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_544.wav", "doc_id": "dvGkKzmIaN.seg_544", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dank.", "score": 100.0}933{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_588.wav", "doc_id": "rISrKoXQCx.seg_588", "src_text": "Thank you for your time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "für deine Zeit.", "score": 91.0}934{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_748.wav", "doc_id": "XejEJmgUmE.seg_748", "src_text": "So that is what we call as the mismatch scenario.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Das ist das sogenannte Missmatch-Szenario.", "score": 96.0}935{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_257.wav", "doc_id": "oYCKgTzTDy.seg_257", "src_text": "To sum up, we build XSemPLR, a unified benchmark for cross-lingual semantic parsing with multiple natural languages and meaning representations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Zusammenfassend können wir sagen, dass wir ein Beispiel für eine einheitliche Referenz für die semantische Analyse mit mehreren natürlichen Sprachen und vielen Repräsentationen erstellen.", "score": 76.0}936{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_517.wav", "doc_id": "dvGkKzmIaN.seg_517", "src_text": "However, this method either not applicable to embedding as services or lack of transferability.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese Methoden sind jedoch entweder nicht anwendbar für die Einbettung von Adressdiensten oder es fehlt an der Übertragbarkeit.", "score": 70.0}937{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_544.wav", "doc_id": "dvGkKzmIaN.seg_544", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dank.", "score": 0.0}938{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_197.wav", "doc_id": "SLpqvupgvW.seg_197", "src_text": "For songs, we simply show a Google search link to each song and then ask the annotators to listen to at least some of each song, and read about each song.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In einigen Fällen zeigen wir einfach einen Google-Suchlink zu jeder Liedes und bitten die Annotatoren, zumindest einige davon zu hören.", "score": 91.0}939{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_176.wav", "doc_id": "SLpqvupgvW.seg_176", "src_text": "The cartoon has three speech bubbles.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Der Cartoon hat drei Sprachblasen.", "score": 100.0}940{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_690.wav", "doc_id": "oaOHnMCwad.seg_690", "src_text": "However these works really don't look at comparing end users with the datasets and models themselves, and studying model and data set positionality is increasingly important as NLP tasks become more subjective and socially oriented, and it's challenging to characterise how these positionalities are skewed because not all decisions are documented and many models are hidden behind APIs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Diese Arbeiten schauen jedoch nicht wirklich darauf, Endnutzer mit den Datensätzen und Modellen selbst zu vergleichen. Das Studieren des Modells und der Positionierbarkeit ist zunehmend wichtig, um mehr subjektiv und sozial orientiert zu sein. Es ist herausfordernd, diese Positionierungen zu beschreiben, da nicht alle Entscheidungen dokumentiert sind und viele Modelle hinter APIs versteckt sind.", "score": 97.0}941{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_58.wav", "doc_id": "TVCREhgqUP.seg_58", "src_text": "Naive seq2seq models struggle with this kind of out-of-distribution generalization and often produce outputs that are detached from the input.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Naive sequenz-zu-Sequenz-Modelle haben Schwierigkeiten mit dieser Art der Verallgemeinerung außerhalb der Verteilung und produzieren oft Ausgaben, die vom Eingabedaten getrennt sind.", "score": 100.0}942{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_463.wav", "doc_id": "SUkmfOTvGi.seg_463", "src_text": "Hello everyone, my name is Shuheng.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hallo alle, mein Name ist Shuhung.", "score": 90.0}943{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_483.wav", "doc_id": "SUkmfOTvGi.seg_483", "src_text": "Here we also found that more fine tuning examples, actually also leads to better generalization.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hier haben wir auch festgestellt, dass mehr Feinabstimmungsbeispiele tatsächlich auch zu einer besseren Verallgemeinerung führen.", "score": 6.0}944{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_396.wav", "doc_id": "WBLMIsdIrq.seg_396", "src_text": "And second, how well do models handle these cases?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und zweitens, wie können die Modelle diese Fälle gut handhaben?", "score": 85.0}945{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_164.wav", "doc_id": "SLpqvupgvW.seg_164", "src_text": "\"Did you mean 'Easy on Me' or 'I Gotta Feeling'?\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Frage: Meinten Sie 'Easy on me' oder 'Ich habe ein Gefühl'?", "score": 100.0}946{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_183.wav", "doc_id": "SLpqvupgvW.seg_183", "src_text": "The first speech bubble is chosen from a few manual prompts per domain.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die erste Sprachblase wird aus ein paar manuellen Prompt pro Domäne ausgewählt.", "score": 65.0}947{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_768.wav", "doc_id": "XejEJmgUmE.seg_768", "src_text": "And the MPP evaluation the way that we do it currently with short and single sentence input, may not fully capture the language models abstract knowledge throughout the context window.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und die MP-Beurteilung, die Art und Weise, wie wir es korrekt mit kurzen und einzelnen Satz-Eingaben durchführen, kann das abstrakte Wissen der Sprachmodelle im Kontextfenster möglicherweise", "score": 60.0}948{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_45.wav", "doc_id": "aQpIWggfCo.seg_45", "src_text": "Thanks for your time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Vielen Dank für Ihre Zeit.", "score": 100.0}949{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_1.wav", "doc_id": "aQpIWggfCo.seg_1", "src_text": "I'm here to introduce our work \"Distilling Script Knowledge from Large Language Models for Constrained Language Planning\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ich bin hier, um unsere Arbeit vorzustellen, die die Unterscheidung von Skriptkenntnissen von Light-Language-Modellen für eingeschränkte Sprachplanung beinhaltet.", "score": 60.0}950{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_734.wav", "doc_id": "XejEJmgUmE.seg_734", "src_text": "Which can also include grammaticality like BLiMP, SyntaxGym, or acceptability in terms of stereotypes such as CrowS pairs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Grammatikalität wie Blimp-Syntakt-Gem oder Akzeptabilität in Form von Stereotypen wie Cruass-Paare umfassen können.", "score": 50.0}951{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_199.wav", "doc_id": "SLpqvupgvW.seg_199", "src_text": "For the recipes and books domain, we show some background text from Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Rezepte und Bücher zeigen wir etwas Hintergrundtext von Wikipedia.", "score": 85.0}952{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_294.wav", "doc_id": "PIZEXUFLAR.seg_294", "src_text": "As we can see, instruction tuning can significantly improve OFA's performance on seen multi-modal tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wie wir sehen können, kann die Anweisungseinstellung die Leistung von O.S. auf Multimode-Aufgaben erheblich verbessern. Auch das", "score": 50.0}953{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_419.wav", "doc_id": "WBLMIsdIrq.seg_419", "src_text": "And finally, we use our benchmark as well as other metrics to evaluate different models on the document-level machine translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und schließlich verwenden wir unsere Benchmarks und andere Metriken, um verschiedene Modelle auf der Dokumentenebene zu bewerten.", "score": 95.0}954{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_78.wav", "doc_id": "TVCREhgqUP.seg_78", "src_text": "We determine the third token in the output in a similar way by jumping to another multiset token.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir bestimmen den dritten Token in der Ausgabe auf ähnliche Weise, indem wir zu einem anderen Multiset-Token springen.", "score": 100.0}955{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_326.wav", "doc_id": "dJGfOSFgZO.seg_326", "src_text": "From our analysis of these evaluation results, we found that ABC-Eval behavior labels are overall more reliable than labels collected by existing methods, as measured by inter-annotator agreement on 100 doubly-labeled conversations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Aus den Ergebnissen dieser Auswertungen geht hervor, dass die A-B-E-Verhaltensetiketten insgesamt zuverlässiger sind als die Etiketten, die durch bestehende Methoden gesammelt wurden. Darüber hinaus sind A-B-C-E-Labels bei der Gesamtkonversationsqualität aussagekräftiger als Metriken, die von existierenden Methoden abgeleitet werden,", "score": 80.0}956{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_788.wav", "doc_id": "WTTtiRKFZI.seg_788", "src_text": "So in English, as you might know, direct objects prefer to be close to the verb, while adjuncts may be further away.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist es also in englischer Sprache, wie du vielleicht weißt, so, dass ein direktes Objekt bevorzugt wird, wenn es der Verfremdung unterworfen ist, während ein Adjunkt vielleicht weiter weg ist, richtig,", "score": 45.0}957{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_784.wav", "doc_id": "WTTtiRKFZI.seg_784", "src_text": "Here loves to all conjuncts separately: Lisa, Bart, and Maggie.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Konjunktionen separat liefern.", "score": 90.0}958{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_431.wav", "doc_id": "hgIDlKNiFM.seg_431", "src_text": "Hi, I am Yanis Labrak and I will present you our works on \"DrBERT: A Robust Pre-trained Model in French for Biomedical and Clinical Domains.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo, ich bin Yan Slavac und ich werde Ihnen unsere Arbeiten über Dr. Bert, ein robustes Trainingsmodell in Französisch für biomedizinische und klinische Bereiche, vorstellen.", "score": 65.0}959{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_229.wav", "doc_id": "oYCKgTzTDy.seg_229", "src_text": "The first one is Translate-Test.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "erste ist der Übersetzungs-Test:", "score": 95.0}960{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_879.wav", "doc_id": "GvEBWkLmuI.seg_879", "src_text": "Have a good time at ACL.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "mir zuhören, haben wir eine gute Zeit in Ägypten.", "score": 30.0}961{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_84.wav", "doc_id": "TVCREhgqUP.seg_84", "src_text": "First of all, the alignment between input and output is not given in the training data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "gelöst. Zunächst ist die Ausrichtung zwischen Eingabe und Ausgabe in den Trainingsdaten nicht angegeben.", "score": 85.0}962{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_718.wav", "doc_id": "oaOHnMCwad.seg_718", "src_text": "So we have a few recommendations for this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "es eine Position in LED und LP gibt? Wir haben also einige Empfehlungen", "score": 90.0}963{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_148.wav", "doc_id": "wLqFAuDnKa.seg_148", "src_text": "And their results so a better performance when using the dev data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Leistung beim Einsatz der Daten ermöglichen.", "score": 85.0}964{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_57.wav", "doc_id": "TVCREhgqUP.seg_57", "src_text": "In this example, the model has seen shallow recursion during training and is tested on an example with deeper recursion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "In diesem Beispiel hat das Modell während des Trainings eine flache Rezursion gesehen und wurde auf einem Beispiel mit einer tiefen Rezursion getestet.", "score": 100.0}965{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_259.wav", "doc_id": "oYCKgTzTDy.seg_259", "src_text": "And our results show many interesting findings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "und unsere Ergebnisse zeigen viele interessante Ergebnisse", "score": 98.0}966{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_64.wav", "doc_id": "TVCREhgqUP.seg_64", "src_text": "Typically, this involves considerable formalism-specific pre-processing of the logical forms, for example, to handle variable symbols.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Typischerweise beinhaltet dies eine vorläufige Verarbeitung der logischen Formen, um z. B. variable Symbole zu handhaben.", "score": 100.0}967{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_360.wav", "doc_id": "gGbuDbHhyc.seg_360", "src_text": "Otherwise, there is a large performance drop.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Andernfalls gibt es einen großen Leistungsverlust,", "score": 100.0}968{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_472.wav", "doc_id": "SUkmfOTvGi.seg_472", "src_text": "This is a data set that we collected from Reuters News from 2020, and then annotated them with the same CoNLL-2003 annotation guidelines.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "+ + +, das wir aus den Nachrichten von Reuters aus dem Jahr 2020 gesammelt und dann mit den gleichen Anmerkungshinweisen Carneal 2003 annotiert haben.", "score": 71.0}969{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_37.wav", "doc_id": "aQpIWggfCo.seg_37", "src_text": "This figure shows the constraint distribution of CoScript.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese Abbildung zeigt die eingeschränkte Verteilung von Coscript.", "score": 100.0}970{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_129.wav", "doc_id": "wLqFAuDnKa.seg_129", "src_text": "This involves using the latest test sets to avoid an overlap of the test data with the training data of the language model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "um eine Überlagerung der Testdaten mit den Trainingsdaten der Sprachmodelle zu vermeiden.", "score": 100.0}971{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_181.wav", "doc_id": "SLpqvupgvW.seg_181", "src_text": "And in the third speech bubble, Bob uses an indirect reference to select one of these entities, for example, \"the newer one.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Frage, und in der dritten Sprechblase wählt Bob eine direkte Referenz aus, um beispielsweise den Neueren auszuwählen. Wir stellen die ersten", "score": 60.0}972{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_121.wav", "doc_id": "uZBWfYjYnf.seg_121", "src_text": "Thanks for your attention.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank für Ihre Aufmerksamkeit.", "score": 100.0}973{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_708.wav", "doc_id": "oaOHnMCwad.seg_708", "src_text": "We find that there is positionality in NLP.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir finden, dass es Positionalität in NLP gibt.", "score": 100.0}974{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_753.wav", "doc_id": "XejEJmgUmE.seg_753", "src_text": "So how does the model do?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wie funktioniert das Modell?", "score": 85.0}975{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_0.wav", "doc_id": "aQpIWggfCo.seg_0", "src_text": "Hi, I'm Siyu Yuan from Fudan University.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo, ich heiße Si Yuyan und komme von der Fudan-Universität.", "score": 60.0}976{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_655.wav", "doc_id": "FLkGnzVRew.seg_655", "src_text": "We find that on transferring the zero-shot performance on the annotated data set is already much better than chance with the best, with AUC .62.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir finden, dass die Übertragung der Null-Schnitt-Leistung auf den annotierten Datensatz bereits viel besser ist als mit dem besten", "score": 60.0}977{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_760.wav", "doc_id": "XejEJmgUmE.seg_760", "src_text": "But when we match the structure, that is when we choose the sentences from the same phenomena in BLiMP or SyntaxGym, we see a massive increase or a massive decrease of the MPP judgement for the model, depending on whether the chosen prefix is acceptable or unacceptable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Aber wenn wir die Struktur wählen, das heißt, wenn wir die Sätze aus demselben Phänomen im Blame-Per-Satz-Grammatik wählen, dann ist das die richtige Struktur. Wir sehen einen massiven Anstieg oder eine massive Abnahme der Einschätzung des Modells durch das Parlament, abhängig davon, ob der gewählte Präfix akzeptabel oder nicht akzeptabel ist.", "score": 60.0}978{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_151.wav", "doc_id": "wLqFAuDnKa.seg_151", "src_text": "In our case, we chose to evaluate with Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "nahe einem kommerziellen System, weshalb wir es mit Google Translate betreiben.", "score": 60.0}979{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_273.wav", "doc_id": "PIZEXUFLAR.seg_273", "src_text": "These tasks are derived from 21 existing open-source dataset and each task is equipped with five expert written instructions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese Aufgaben sind von einundzwanzig vorhandenen Open-Source-Datensätzen abgeleitet, und jede Aufgabe ist mit fünf zusätzlichen Anweisungen ausgestattet.", "score": 70.0}980{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_32.wav", "doc_id": "aQpIWggfCo.seg_32", "src_text": "However, previous studies do not enable planning for specific goals and manual dataset annotation is expensive.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "dorthin. Vorherige Studien ermöglichen jedoch keine Planung für spezifische Ziele, und die manuelle Datensatzannotation ist aufwendig.", "score": 95.0}981{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_643.wav", "doc_id": "FLkGnzVRew.seg_643", "src_text": "Studying dissonance expressed in language can also be beneficial in understanding extremism and polarization of vulnerable groups.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Untersuchung der Abstände in der Sprache kann auch für das Verständnis von Extremismus und Polarisierung von Gruppen sinnvoll sein.", "score": 60.0}982{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_67.wav", "doc_id": "TVCREhgqUP.seg_67", "src_text": "For the first time, we show strong generalization to deeper recursion without relying on trees.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zum ersten Mal sehen wir eine starke Verallgemeinerung, um die Rekonstruktion durchzuführen, ohne auf Tricks zurückzugreifen.", "score": 35.0}983{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_703.wav", "doc_id": "oaOHnMCwad.seg_703", "src_text": "We've then compared these, annotations with Social Chemistry, Delphi and GPT 4.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir verglichen diese Anmerkungen dann mit Social Chemistry, Delphi und GPT-4. Wir werden", "score": 60.0}984{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_219.wav", "doc_id": "oYCKgTzTDy.seg_219", "src_text": "As shown in this figure, we need to translate the query in multiple natural languages using neural models to SQL, Lambda or FunQL, and etcetera.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wie in diesem Bild gezeigt, müssen wir die Anfrage in mehrere natürliche Sprachen übersetzen, indem wir neuere Modelle verwenden: 2, Ceql, Lmda oder FQL und", "score": 49.0}985{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_650.wav", "doc_id": "FLkGnzVRew.seg_650", "src_text": "To no surprise, the classifier performed not much better than chance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "von Unterschieden, um keine Überraschung zu erzeugen, dass die Klassifikation nicht viel besser ist als die Chance.", "score": 30.0}986{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_805.wav", "doc_id": "WTTtiRKFZI.seg_805", "src_text": "Right?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Also, was", "score": 0.0}987{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_768.wav", "doc_id": "XejEJmgUmE.seg_768", "src_text": "And the MPP evaluation the way that we do it currently with short and single sentence input, may not fully capture the language models abstract knowledge throughout the context window.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "sind, und dass die MPPE-Bewertung, die wir derzeit mit kurzen und einzelnen Sätzen als Eingabe durchführen, möglicherweise nicht vollständig die abstrakte Sprachmodelle-Kenntnisse im Kontextfenster erfassen.", "score": 99.0}988{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_233.wav", "doc_id": "oYCKgTzTDy.seg_233", "src_text": "In this setting, the source language is the same as target language, for example German to German or English to English.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In dieser Konstellation ist die Quellsprache dieselbe wie die Zielsprache, zum Beispiel Deutsch zu Deutsch oder Englisch zu Englisch.", "score": 99.0}989{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_372.wav", "doc_id": "gGbuDbHhyc.seg_372", "src_text": "To summarize, we showed that recent WSL approaches require clean, manually annotated samples for them to work properly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Zusammengefasst zeigen wir, dass aktuelle WSL-Ansätze saubere, manuell annotierte Proben erfordern, damit sie ordnungsgemäß funktionieren.", "score": 100.0}990{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_284.wav", "doc_id": "PIZEXUFLAR.seg_284", "src_text": "So we use pre-trained OFA large model as a base model.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Daher verwenden wir ein vorbereitetes OFA-Modell als Basismodell;", "score": 95.0}991{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_722.wav", "doc_id": "oaOHnMCwad.seg_722", "src_text": "And a good example of this is the Masakhani initiative.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und ein gutes Beispiel dafür", "score": 100.0}992{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_123.wav", "doc_id": "wLqFAuDnKa.seg_123", "src_text": "This is joint work with my colleagues from Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies ist eine Zusammenarbeit mit meinen Kollegen von Google Translate.", "score": 100.0}993{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_142.wav", "doc_id": "wLqFAuDnKa.seg_142", "src_text": "And when we go, as in our case, to five-shot prompting, there is nearly no difference to the actual form of the prompting.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Schuss, und wenn wir, wie in unserem Fall, zum Fächerschießen gehen, gibt es kaum einen Unterschied zur tatsächlichen Form des Schießens. Es", "score": 67.0}994{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_421.wav", "doc_id": "WBLMIsdIrq.seg_421", "src_text": "But then if we use COMET, context-aware models perform best.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die beste Leistung aufweisen. Aber", "score": 67.0}995{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_163.wav", "doc_id": "SLpqvupgvW.seg_163", "src_text": "Consider this alternative question.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Betrachten Sie diese alternative", "score": 95.0}996{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_626.wav", "doc_id": "oeooqChmKK.seg_626", "src_text": "Additional experiments with fictional knowledge indicated even the best performing models, cannot reliably integrate backward knowledge provided only at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zusätzliche Experimente mit fiktiverem Wissen zeigen, dass selbst die besten Leistungsmodelle nicht zuverlässig Hintergrundwissen, das nur zur Zeit der Erinnerung angeboten wird, integrieren können.", "score": 95.0}997{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_595.wav", "doc_id": "oeooqChmKK.seg_595", "src_text": "Pretrained parameters can contain information about what presidents do and what a TV is but they cannot reliably know who this instance-specific entity \"John\" is, or who the new president is, because the president might have changed since pretraining.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "können Informationen über das, was Präsidenten tun und was sie sind, enthalten, aber sie können nicht zuverlässig wissen, wer diese spezifische Einheit ist oder wer der neue Präsident ist, weil der Präsident sich vielleicht während des Prä-Trainings verändert", "score": 95.0}998{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_777.wav", "doc_id": "WTTtiRKFZI.seg_777", "src_text": "Right.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "beiden", "score": 0.0}999{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_42.wav", "doc_id": "aQpIWggfCo.seg_42", "src_text": "We evaluate constrained language planning ability of large language models and develop an over-generate-then-filter method for large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "wir bewerten die konstrizierten Sprachplanungsfähigkeiten von großen Sprachmodellen und entwickeln ein übergenerierendes Filterverfahren für große Sprachmodelle.", "score": 98.0}1000{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_100.wav", "doc_id": "uZBWfYjYnf.seg_100", "src_text": "So what is our solution?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "ist unsere Lösung?", "score": 100.0}1001{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_644.wav", "doc_id": "FLkGnzVRew.seg_644", "src_text": "Finally, cognitive dissonance is important to understand personal cognitive styles of individuals and helps us understand decision making processes better.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Schließlich ist es wichtig, kontextuelle Unterschiede zu verstehen, um persönliche kontextuelle Stile von Einzelpersonen zu verstehen und uns zu helfen, Entscheidungsprozesse besser zu verstehen.", "score": 90.0}1002{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_128.wav", "doc_id": "wLqFAuDnKa.seg_128", "src_text": "We evaluated the transition capability of such models using the best practices of the MT community.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir bewerten die Übersetzbarkeit von Modellen, indem wir die besten Übersetzungen der Gemeinschaft verwenden,", "score": 55.0}1003{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_182.wav", "doc_id": "SLpqvupgvW.seg_182", "src_text": "We provide the first and second speech bubbles automatically, but the third one is filled in by the annotator.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir liefern die ersten und zweiten Sprechblasen automatisch, aber die dritte wird vom Annotator eingegeben.", "score": 81.0}1004{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_741.wav", "doc_id": "XejEJmgUmE.seg_741", "src_text": "So that is the approach.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Also ist das der Ansatz,", "score": 99.0}1005{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_806.wav", "doc_id": "WTTtiRKFZI.seg_806", "src_text": "It violates one principle, but it satisfies another one.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Es verletzt ein Prinzip, aber es erfüllt ein anderes.", "score": 100.0}1006{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_302.wav", "doc_id": "PIZEXUFLAR.seg_302", "src_text": "We also can see transfer learning from natural instruction datasets can help OFA to attain much better performance on the natural instruct dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir können auch sehen, dass das Transfer-Lernen aus dem Datensatz der natürlichen Anweisung die Leistung von OFA auf dem Datensatz der natürlichen Anweisung verbessern kann.", "score": 95.0}1007{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_29.wav", "doc_id": "aQpIWggfCo.seg_29", "src_text": "Our method greatly improves the planning ability both in semantic completeness and faithfulness to the constraint.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unsere Methode verbessert die Planbarkeit erheblich sowohl in Bezug auf semantische Vollständigkeit als auch in Bezug auf Treue zu den Einschränkungen.", "score": 100.0}1008{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_712.wav", "doc_id": "oaOHnMCwad.seg_712", "src_text": "We also find most additional alignment with people who have a college education.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir finden auch heraus, dass die meisten zusätzlichen Übereinstimmungen mit Personen mit einem College-Abschluss bestehen,", "score": 99.0}1009{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_46.wav", "doc_id": "aQpIWggfCo.seg_46", "src_text": "Please find more details of CoScript in our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Bitte finden Sie weitere Einzelheiten zu Co-Script in unseren Unterlagen.", "score": 95.0}1010{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_31.wav", "doc_id": "aQpIWggfCo.seg_31", "src_text": "Creating the dataset is an essential step to this end.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Das Erstellen eines Datensatzes ist ein wesentlicher Schritt auf dem Weg", "score": 100.0}1011{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_782.wav", "doc_id": "WTTtiRKFZI.seg_782", "src_text": "And finally, there's also a multi-headed approach that's used, for example, in the Hudson's Word Grammar, where they say all conjuncts are heads of the coordinate structure.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und schließlich ist dies auch ein mehrfach überarbeiteter Ansatz, der beispielsweise in der Katschons-Word-grammatik verwendet wird. Wo, so zu sagen, alle Konjungate über die Koordinatenstruktur stehen,", "score": 67.0}1012{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_776.wav", "doc_id": "WTTtiRKFZI.seg_776", "src_text": "So these two approaches are asymmetric.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Ansätze gleich sind. Nun", "score": 67.0}1013{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_268.wav", "doc_id": "PIZEXUFLAR.seg_268", "src_text": "Additionally, at the time of our research, we discovered a considerable discrepancy in the availability of instructional datasets between NLP and multi-modal.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Darüber hinaus entdeckten wir in der Zeit unserer Forschung eine erhebliche Diskrepanz in der Verfügbarkeit von Anweisungsdatensätzen zwischen L und M.", "score": 60.0}1014{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_591.wav", "doc_id": "oeooqChmKK.seg_591", "src_text": "Natural language understanding models draw on a variety of knowledge sources, such as knowledge contained in their parameters, usually acquired by a pretraining, and knowledge given in inputs at inference time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Modelle zur Verständigung von nationalen Sprachen basieren auf einer Vielzahl von Wissensquellen, wie z. B. Wissen, das in ihren Parametern enthalten ist, das normalerweise durch eine Vorbereitung erworben wird, und Wissen, das in Eingaben bei der Zeit der", "score": 90.0}1015{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_123.wav", "doc_id": "wLqFAuDnKa.seg_123", "src_text": "This is joint work with my colleagues from Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies ist eine gemeinsame Arbeit mit meinen Kollegen von Google Translate.", "score": 100.0}1016{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_467.wav", "doc_id": "SUkmfOTvGi.seg_467", "src_text": "We observe that models have been used in CoNLL-2003 to develop NER for almost 20 years and this naturally raises several problems.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir haben festgestellt, dass Modelle seit fast zwanzig Jahren für die Entwicklung von Neuronen in Corel Draw verwendet werden, und dies wirft natürlich einige Probleme auf.", "score": 3.0}1017{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_506.wav", "doc_id": "dvGkKzmIaN.seg_506", "src_text": "Embedding as services is one of the services built upon large language models to assist various, NLP tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "ist einer der Dienste, die auf großen Sprachmodellen basieren, um verschiedene NLP-Aufgaben zu unterstützen.", "score": 100.0}1018{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_823.wav", "doc_id": "WTTtiRKFZI.seg_823", "src_text": "But when the governor is on the right this tendency disappears.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "auf der rechten Seite beobachtet.", "score": 10.0}1019{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_565.wav", "doc_id": "rISrKoXQCx.seg_565", "src_text": "And we also try to investigate whether language models can pick up the polarisation that's prevalent in our modern society.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und wir werden auch versuchen, die Polarisierung, die in unserer modernen Gesellschaft vorherrscht, zu untersuchen.", "score": 1.0}1020{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_251.wav", "doc_id": "oYCKgTzTDy.seg_251", "src_text": "The orange line is Cross-lingual Zero-shot transfer.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "orangefarbene Linie ist der Cross-Lingual-Zero-Shot-Transfer, während die", "score": 100.0}1021{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_61.wav", "doc_id": "TVCREhgqUP.seg_61", "src_text": "The trees are intended to capture the compositional process that relates utterances with the logical forms.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Bäume sollen das kompositionelle Verfahren erfassen, das sich mit den logischen Formen in Verbindung setzt.", "score": 95.0}1022{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_822.wav", "doc_id": "WTTtiRKFZI.seg_822", "src_text": "What we see here is that when the governor is on the left, the tendency for the left conjunct to be shorter grows steadily, with the absolute difference in words, and the same is observed when there is no governor as in coordination of sentences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Was wir hier sehen, ist, dass, wenn die Regierung Die Tendenz, dass die linke Konjunktion kürzer ist, wächst stetig mit der absoluten Differenz in Wörtern und dasselbe wird beobachtet, wenn es keine Gouverneur gibt, aber wenn der Gouverneur", "score": 56.0}1023{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_404.wav", "doc_id": "WBLMIsdIrq.seg_404", "src_text": "We perform our analysis at three different levels.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir führen unsere Analysen auf drei verschiedenen Ebenen durch.", "score": 100.0}1024{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_190.wav", "doc_id": "SLpqvupgvW.seg_190", "src_text": "The first one is uniform at random.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die erste ist Uniform-Attrack", "score": 75.0}1025{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_4.wav", "doc_id": "aQpIWggfCo.seg_4", "src_text": "And show that large language models can effectively decompose goals into steps.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und gezeigt, dass große Sprachmodelle Ziele effektiv in Schritte zerlegen können.", "score": 100.0}1026{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_569.wav", "doc_id": "rISrKoXQCx.seg_569", "src_text": "So this indicates that language models can also pick up the polarisation in our society.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "im Zentrum befinden, so dass die Sprachmodelle auch die Polarisierung in unserer Gesellschaft abbilden können.", "score": 95.0}1027{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_829.wav", "doc_id": "GvEBWkLmuI.seg_829", "src_text": "This work is done in collaboration with Esin Durmus and Dan Jurafsky.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "zu messen. Diese Arbeit wird in Zusammenarbeit mit Esnader Mush und Danarovsky durchgeführt.", "score": 68.0}1028{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_873.wav", "doc_id": "GvEBWkLmuI.seg_873", "src_text": "So based on these patterns, we conclude with three recommendations for model owners.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Auf der Grundlage dieser Muster können wir drei Empfehlungen für Modelleigentümer zusammenfassen.", "score": 95.0}1029{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_72.wav", "doc_id": "TVCREhgqUP.seg_72", "src_text": "We introduce a new method to predict the permutation that does not put any hard constraints on the possible permutations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir führen eine neue Methode zur Vorhersage der Permutation ein, die keine harten Einschränkungen für die möglichen Permutationen", "score": 99.0}1030{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_438.wav", "doc_id": "hgIDlKNiFM.seg_438", "src_text": "Since its release in 2018, BERT has become one of the most effective approach to solve natural language processing tasks and offers huge performance gains compared to historical static and contextualized methods such as Word2vec, fastText, or more.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Seit seiner Veröffentlichung im Jahr 2018 ist Bert ein effektiver Ansatz zur Lösung natürlicher Sprachverarbeitungsaufgaben und bietet im Vergleich zu historischen statischen und kontextualisierten Methoden wie Word-to-Vec, FastText oder ANW einen größeren Leistungsgewinn.", "score": 99.0}1031{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_256.wav", "doc_id": "oYCKgTzTDy.seg_256", "src_text": "Pretraining on English natural language can significantly boost the performance of Few-shot on target natural languages, and we found multilingual language models such as Codex and BLOOM are still inadequate for cross-lingual semantic parsing tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Das Training auf Englisch kann die Leistung von Few-Shot auf Ziel-Natursprachen erheblich verbessern und wir fanden heraus, dass multilinguale Sprachmodelle wie Coders und Blue immer noch unzureichend für die Überprüfung von Semantik in mehreren Sprachen sind.", "score": 95.0}1032{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_784.wav", "doc_id": "WTTtiRKFZI.seg_784", "src_text": "Here loves to all conjuncts separately: Lisa, Bart, and Maggie.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Gouverneur, hier liebt, zu allen Konjunktionen separat, diese sind aber nicht relevant.", "score": 0.0}1033{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_317.wav", "doc_id": "dJGfOSFgZO.seg_317", "src_text": "However, we believe there is a more precise and reliable strategy for dimensional dialogue evaluation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir glauben jedoch, dass es eine genauer und zuverlässigere Strategie für die dimensionale Dialogbewertung gibt.", "score": 97.0}1034{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_671.wav", "doc_id": "FLkGnzVRew.seg_671", "src_text": "These are the links to our core data set and our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Diese sind die Links zu unserem Code-Datensatz und unserem Papier.", "score": 95.0}1035{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_535.wav", "doc_id": "dvGkKzmIaN.seg_535", "src_text": "We compute the similarity difference between benign and backdoor data set which is defined as delta cosine and delta L2.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir berechnen den Ähnlichkeitsunterschied zwischen dem Basis-Embedding und dem Backdoor-Embedding, das als Delta-Embedding definiert ist.", "score": 90.0}1036{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_131.wav", "doc_id": "wLqFAuDnKa.seg_131", "src_text": "We use state-of-the-art, neural MT metrics, and additionally also show expert-based human evaluation results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir verwenden state-of-the-art-neuronale MT-Metriken und zeigen zusätzlich Ergebnisse der Expertenbewertung durch Menschen.", "score": 90.0}1037{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_282.wav", "doc_id": "PIZEXUFLAR.seg_282", "src_text": "We use all the instances in the test split for each task.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir verwenden alle Instanzen im Test für", "score": 80.0}1038{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_668.wav", "doc_id": "FLkGnzVRew.seg_668", "src_text": "However, the annotators also find the examples difficult.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Anmerkungen sind jedoch schwierig. In der Zusammenfassung", "score": 0.0}1039{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_830.wav", "doc_id": "GvEBWkLmuI.seg_830", "src_text": "In recent years, many have documented the prevalence of social bias and stereotypes in large language models, or LLMs.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In jüngster Zeit haben viele die Prävalenz von sozialen Vorurteilen und Stereotypen in großen Sprachmodellen dokumentiert.", "score": 100.0}1040{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_275.wav", "doc_id": "PIZEXUFLAR.seg_275", "src_text": "OFA uses a unified vocabulary for language, image tokens and the coordinates of a bounding box.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "verwendet ein einheitliches Vokabular für Sprache, Bildsymbole und Koordinaten von Begrenzungskästen.", "score": 100.0}1041{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_49.wav", "doc_id": "TVCREhgqUP.seg_49", "src_text": "This is joint work with my advisors Alexander Koller and Ivan Titov.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies ist eine gemeinsame Arbeit mit meinen Beratern Alexander Koller und Ivan Titov.", "score": 100.0}1042{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_669.wav", "doc_id": "FLkGnzVRew.seg_669", "src_text": "In summary, we find that PRC is a simple AL strategy for rare class acquisition and cold starting AL with appropriately designed transfer learning task and help significantly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zusammenfassend finden wir, dass PRC eine einfache AL-Strategie für die Akquisition von Raritätsklassen ist und die Ko-Start-AL mit entsprechend konzipierten Transfer-Lernaufgaben erheblich unterstützen kann.", "score": 90.0}1043{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_6.wav", "doc_id": "aQpIWggfCo.seg_6", "src_text": "Planning for the goals with specific constraints, such as \"make a chocolate cake\", still remains under-studied.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Planung von Zielen mit spezifischen Einschränkungen, wie z. B. Make a Chocolate Cake, ist noch nicht untersucht.", "score": 40.0}1044{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_322.wav", "doc_id": "dJGfOSFgZO.seg_322", "src_text": "For example, ABC-Eval measures the number of turns in which a chat model ignores its partner or says something irrelevant, contradicts itself or its partner, hallucinates incorrect facts or violates common sense knowledge, and when the model succeeds or fails to show empathy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "A. C. E. die Anzahl der Umdrehungen, die ein Chat-Modell ignoriert oder relevant ist. Es widerspricht sich selbst oder seinem Partner, halluziniert unkorrekte Fakten oder verletzt das allgemeine Menschenverständnis, und wenn das Modell erfolgreich ist oder nicht, zeigt es Empathie.", "score": 60.0}1045{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_493.wav", "doc_id": "SUkmfOTvGi.seg_493", "src_text": "And these goes hand in hand, we can't just have one ingredient but throw out the others.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Beispiele benötigen würden. Zur gleichen Zeit stellten", "score": 0.0}1046{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_16.wav", "doc_id": "aQpIWggfCo.seg_16", "src_text": "Then we conduct detailed analysis to investigate why learning models fail.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "führen wir detaillierte Analysen durch, um zu untersuchen, was die linearen Modelle für die Ergebnisse verantwortlich sind. Die Ergebnisse", "score": 80.0}1047{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_75.wav", "doc_id": "TVCREhgqUP.seg_75", "src_text": "We go from left to right over the output and determine which multiset token to put in every position.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir gehen von links nach rechts über die Ausgabe und bestimmen, welcher Multiset-Token in jede Position gesetzt werden soll.", "score": 100.0}1048{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_810.wav", "doc_id": "WTTtiRKFZI.seg_810", "src_text": "And, also the observation that was made in parsing that this tendency grows with length difference.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Und auch die Beobachtung, dass das Wachstum der Tendenz mit Längenunterschieden einherging. Wenn", "score": 90.0}1049{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_751.wav", "doc_id": "XejEJmgUmE.seg_751", "src_text": "Finally, we can choose sentences from a completely unrelated domain such as Wikipedia.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Schließlich können wir Sätze aus einem vollkommen unabhängigen Bereich wie Wikipedia auswählen.", "score": 100.0}1050{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_244.wav", "doc_id": "oYCKgTzTDy.seg_244", "src_text": "We found that Encoder-Decoder obtains the best performance on all nine datasets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir fanden heraus, dass Encoder-Decoder-Modelle bessere Ergebnisse erzielen als monolinguale Modelle. Wir bewerten", "score": 85.0}1051{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_603.wav", "doc_id": "oeooqChmKK.seg_603", "src_text": "Servin and Kea met at a park.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Bäckerin. Serwin und Kiah trafen sich nach einem", "score": 40.0}1052{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_429.wav", "doc_id": "WBLMIsdIrq.seg_429", "src_text": "Thank you so much for your attention.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Vielen Dank für die Unterstützung.", "score": 80.0}1053{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_628.wav", "doc_id": "oeooqChmKK.seg_628", "src_text": "However, with task-specific training, some models successfully integrate knowledge from multiple sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Mit spezifischem Training können einige Modelle jedoch erfolgreich Wissen aus mehreren Quellen integrieren.", "score": 99.0}1054{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_745.wav", "doc_id": "XejEJmgUmE.seg_745", "src_text": "We extract grammatical sentences from Adjunct Island and then we add it as a prefix to both the acceptable query and the unacceptable query.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir extrahieren grammatikalische Sätze aus dem Adjektiv. Und dann fügen wir es als Präfix sowohl zur akzeptablen als auch zur inakzeptablen Frage hinzu.", "score": 95.0}1055{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_45.wav", "doc_id": "aQpIWggfCo.seg_45", "src_text": "Thanks for your time.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Vielen Dank für Ihre Zeit.", "score": 100.0}1056{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_703.wav", "doc_id": "oaOHnMCwad.seg_703", "src_text": "We've then compared these, annotations with Social Chemistry, Delphi and GPT 4.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dann verglichen wir diese Anmerkungen mit Social Chemistry, Delphy und GPD Four. Dann", "score": 60.0}1057{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_580.wav", "doc_id": "rISrKoXQCx.seg_580", "src_text": "We would also like to highlight that we expose the unique dilemma regarding language model political biases.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir würden also gerne auch darauf hinweisen, dass wir das einzigartige Dilemma, das sich bei der Untersuchung der monologischen politischen Auffassungen", "score": 85.0}1058{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_598.wav", "doc_id": "oeooqChmKK.seg_598", "src_text": "We introduce a coreference resolution task, designed to probe for the ability to draw on knowledge available in different sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir führen eine Korrelationsanalyse durch, die darauf ausgelegt ist, die Fähigkeit zu testen, Wissen aus verschiedenen Quellen abzurufen.", "score": 95.0}1059{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_583.wav", "doc_id": "rISrKoXQCx.seg_583", "src_text": "If we do try to sanitaze somehow, we would also risk censorship, or exclusion.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn wir versuchen, uns zu sänitieren, würden wir auch Risiken der Zensur oder Ausgrenzung laufen, und", "score": 83.0}1060{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_109.wav", "doc_id": "uZBWfYjYnf.seg_109", "src_text": "If we go on and we receive another speech chunk, and our model predicts other three words and we will look at those cross-attention weights, we will see that no word points to the last lambda speech frames.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wenn wir weitergehen, erhalten wir einen weiteren Sprachblock, und unser Modell sagt uns drei weitere Wörter, und wir schauen uns die Wechselwirkung an. Wir werden erkennen, dass kein Wort auf die letzte, lambe, lambe, lambe, lambe, lambe, lambe, lambe, lambe, lambe.", "score": 0.0}1061{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_302.wav", "doc_id": "PIZEXUFLAR.seg_302", "src_text": "We also can see transfer learning from natural instruction datasets can help OFA to attain much better performance on the natural instruct dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir können auch sehen, dass das Übertragen aus den Datensätzen der natürlichen Anweisung OIF helfen kann, um eine viel bessere Leistung auf den Datensätzen der natürlichen Anweisung zu erzielen.", "score": 66.0}1062{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_730.wav", "doc_id": "XejEJmgUmE.seg_730", "src_text": "Language model acceptability judgments are not always robust to context.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die Akzeptanzurteile des Sprachmodells sind nicht immer robust.", "score": 66.0}1063{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_692.wav", "doc_id": "oaOHnMCwad.seg_692", "src_text": "We do this through our framework NLPositionality.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir tun dies durch unser Framework.", "score": 85.0}1064{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_121.wav", "doc_id": "uZBWfYjYnf.seg_121", "src_text": "Thanks for your attention.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank für Ihre Aufmerksamkeit.", "score": 100.0}1065{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_219.wav", "doc_id": "oYCKgTzTDy.seg_219", "src_text": "As shown in this figure, we need to translate the query in multiple natural languages using neural models to SQL, Lambda or FunQL, and etcetera.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wie in der Abbildung zu sehen ist, müssen wir den Query in mehrere natürliche Sprachen übersetzen, indem wir Neuronenmodelle verwenden, wie z. B. Seq2Seq, Lambada oder Funktions-Query.", "score": 41.0}1066{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_385.wav", "doc_id": "WBLMIsdIrq.seg_385", "src_text": "This work was done in collaboration with Patrick Fernandes, Emmy Liu, André F. T. Martins, and Graham Neubig.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Diese Arbeit wurde in Zusammenarbeit mit Patrick Fennan, M.E., und Andrew F. Martens durchgeführt. Die Übersetzungen hängen also", "score": 14.0}1067{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_74.wav", "doc_id": "TVCREhgqUP.seg_74", "src_text": "Conceptually, our permutation model works roughly like this.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Konzeutell, unser Permutationenmodell funktioniert ungefähr so.", "score": 87.0}1068{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_684.wav", "doc_id": "oaOHnMCwad.seg_684", "src_text": "Positionality is simply the perspectives that people hold as a result of their demographics, identity, and life experiences.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "ist einfach die Perspektive, die die Menschen aufgrund ihrer Demografie, Identität und Lebenserfahrungen haben.", "score": 85.0}1069{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_727.wav", "doc_id": "oaOHnMCwad.seg_727", "src_text": "Thank you.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Vielen Dank.", "score": 100.0}1070{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_532.wav", "doc_id": "dvGkKzmIaN.seg_532", "src_text": "Back door data set contains sentences of which all words belong to the trigger set while all words in the sentences of benign data set do not belong to the trigger sets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Datensatz; der Datensatz der Rückwand enthält Sätze, deren alle Wörter dem Trigger-Set gehören, während alle Wörter in den Sätzen des bösartigen Datensatzes nicht dem Trigger-Set gehören.", "score": 88.0}1071{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_66.wav", "doc_id": "TVCREhgqUP.seg_66", "src_text": "In this paper, we don't use trees and introduce a neural seq2seq model that directly models the correspondences between fragments of the input and fragments of the output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In diesem Papier verwenden wir keine Traces und stellen ein neues Sequenz-zu-Sequenz-Modell vor, das die Korrespondenzen zwischen den Fragmenten des Inputs und den Fragmenten des Outputs direkt modelliert.", "score": 100.0}1072{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_773.wav", "doc_id": "WTTtiRKFZI.seg_773", "src_text": "So for example, in the universal dependencies, the structure of the coordination, Lisa, Bart, and Maggie, such that the first conjunct is the head of the whole coordinate structure.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "dass die Struktur der Abhängigkeitskoordination Lisa Bart und Meggie ist. Es ist so, dass der erste Konjunkt der Kopf der ganzen Struktur ist.", "score": 65.0}1073{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_460.wav", "doc_id": "hgIDlKNiFM.seg_460", "src_text": "We are also observing that more specialized data is better, but it doesn't scale well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir stellen auch fest, dass spezialisierte Daten besser sind - mehr spezialisierte Daten sind besser - aber sie werden nicht gut genutzt.", "score": 57.0}1074{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_19.wav", "doc_id": "aQpIWggfCo.seg_19", "src_text": "The heat map in the figure shows that the planning performance of InstructGPTs varies considerably for goals of different categories.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Übersicht im Bild zeigt, dass die Planungsleistung von Mädchen in verschiedenen Kategorien sehr unterschiedlich", "score": 44.0}1075{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_423.wav", "doc_id": "WBLMIsdIrq.seg_423", "src_text": "This again demonstrates that it is difficult to determine the best document-level translation system if we use corpus-level metrics alone.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dies zeigt wieder, dass es schwierig ist, das beste Dokumenten-Übersetzungs-System zu bestimmen, wenn man nur Korpus-Ebenniveaumetriken verwendet. Jetzt", "score": 100.0}1076{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_611.wav", "doc_id": "oeooqChmKK.seg_611", "src_text": "We have defined three settings of KITMUS.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir haben drei Einstellungen für Kidmus definiert.", "score": 95.0}1077{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_205.wav", "doc_id": "SLpqvupgvW.seg_205", "src_text": "The AltEntities Corpus has 6,000 alternative questions across three domains, and it has 42,000 indirect referring expressions.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "tausend alternative Fragen in drei Domänen und zweitausend indirekte Referenzäußerungen.", "score": 50.0}1078{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_337.wav", "doc_id": "dJGfOSFgZO.seg_337", "src_text": "However, this is all the more reason to pursue reliable and precise evaluation metrics for comparing models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Dies ist jedoch der Grund, warum wir zuverlässige und präzise Bewertungsmetriken für Vergleichsmodelle verwenden.", "score": 90.0}1079{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_226.wav", "doc_id": "oYCKgTzTDy.seg_226", "src_text": "We provide a uniform data set XSemPLR for cross-lingual semantic parsing in multiple natural languages and meaning representations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "vor, wir stellen ein einheitliches Datensatz-Beispiel für das Cross-Lingual-Semantic-Parsing in mehreren natürlichen Sprachen und Bedeutungsdarstellungen bereit. Es", "score": 99.0}1080{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_396.wav", "doc_id": "WBLMIsdIrq.seg_396", "src_text": "And second, how well do models handle these cases?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Kontext? Zweitens: Wie gut können Modelle diese Fälle handhaben?", "score": 100.0}1081{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_261.wav", "doc_id": "oYCKgTzTDy.seg_261", "src_text": "And welcome to visit our paper and code.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Aufmerksamkeit.", "score": 0.0}1082{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_10.wav", "doc_id": "aQpIWggfCo.seg_10", "src_text": "In this paper, we first evaluate and improve the constrained language planning ability of large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "In diesem Papier bewerten und verbessern wir zunächst die eingeschränkte Planungsfähigkeit von Sprachmodellen in Großschreibung. Außerdem gibt", "score": 75.0}1083{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_793.wav", "doc_id": "WTTtiRKFZI.seg_793", "src_text": "Because then it can be moved to the position after the adjunct.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "weil es dann in die Position nach dem Add-on bewegt", "score": 95.0}1084{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_496.wav", "doc_id": "SUkmfOTvGi.seg_496", "src_text": "And we found that the answer is actually a resounding yes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wir haben festgestellt, dass die Antwort tatsächlich ein lautes „Ja“ ist.", "score": 95.0}1085{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_78.wav", "doc_id": "TVCREhgqUP.seg_78", "src_text": "We determine the third token in the output in a similar way by jumping to another multiset token.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir bestimmen den dritten Token in der Ausgabe auf ähnliche Weise, indem wir zu einem anderen Multisets-Token springen;", "score": 98.0}1086{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_810.wav", "doc_id": "WTTtiRKFZI.seg_810", "src_text": "And, also the observation that was made in parsing that this tendency grows with length difference.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Und auch die Beobachtung, die gemacht wurde, als sie vorüberging, dass eine Dissonanz mit langen Unterschieden wächst,", "score": 65.0}1087{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_722.wav", "doc_id": "oaOHnMCwad.seg_722", "src_text": "And a good example of this is the Masakhani initiative.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und ein gutes Beispiel dafür", "score": 99.0}1088{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_305.wav", "doc_id": "PIZEXUFLAR.seg_305", "src_text": "So one more thing, we are collecting a much larger multi-model instruction tuning dataset with around 150 additional vision language tasks and we will release them.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "für die Anpassung mit etwa einhundertfünfzig zusätzlichen Sprachaufgaben und geben sie heraus.", "score": 80.0}1089{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_117.wav", "doc_id": "uZBWfYjYnf.seg_117", "src_text": "And we see that it outperforms all the strategies applied to offline models since the curves are shifted over the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und wir sehen, dass alle Strategien, die auf Offline-Modelle angewendet werden, seit den Kurven nach links verschoben sind.", "score": 48.0}1090{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_589.wav", "doc_id": "oeooqChmKK.seg_589", "src_text": "Hello everyone, I'm Akshatha, and today my co-author Martin and I are presenting our work \"The KITMUS Test: Evaluating Knowledge Integration from Multiple Sources.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hallo alle, ich bin Ashutosh und heute präsentieren mein Co-Autor Martin und ich unsere Arbeit, den KITMST-Test: Die Bewertung der Integration von Wissen aus mehreren Quellen.", "score": 53.0}1091{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_258.wav", "doc_id": "oYCKgTzTDy.seg_258", "src_text": "We conduct a comprehensive benchmark study on three representative types of multilingual language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir führen eine umfassende Benchmark-Studie an drei repräsentativen Typen von mehrsprachigen Modellen durch", "score": 100.0}1092{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_512.wav", "doc_id": "dvGkKzmIaN.seg_512", "src_text": "First the method should be applicable to embedding as services.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Erstens sollte die Methode auf eingebettete ET-Verbindungen anwendbar sein.", "score": 52.0}1093{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_203.wav", "doc_id": "SLpqvupgvW.seg_203", "src_text": "Here are some examples from our dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hier sind einige Beispiele aus unserem Datensatz.", "score": 100.0}1094{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_512.wav", "doc_id": "dvGkKzmIaN.seg_512", "src_text": "First the method should be applicable to embedding as services.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Erstens sollte die Methode für die Einbettung von Dienstleistungen anwendbar sein.", "score": 95.0}1095{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_726.wav", "doc_id": "oaOHnMCwad.seg_726", "src_text": "But if you'd like to learn more, feel free to check out our dashboard for the most updated analysis results and our paper.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Empfehlung. Wenn Sie jedoch mehr erfahren möchten, können Sie gerne unser Dashboard für die neuesten Analyseergebnisse und unser Papier überprüfen.", "score": 90.0}1096{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_432.wav", "doc_id": "hgIDlKNiFM.seg_432", "src_text": "In this presentation, we first talk about language modeling in healthcare.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In dieser Präsentation sprechen wir zunächst über Sprachmodellierung in der Gesundheitsversorgung,", "score": 100.0}1097{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_294.wav", "doc_id": "PIZEXUFLAR.seg_294", "src_text": "As we can see, instruction tuning can significantly improve OFA's performance on seen multi-modal tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "die Leistung von Multimodaltasks erheblich verbessern. Auch das Transfer-Lernen aus", "score": 60.0}1098{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_874.wav", "doc_id": "GvEBWkLmuI.seg_874", "src_text": "First, we should, as researchers, be addressing positive stereotypes and essentializing narratives.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Eigentümer von Modellen zusammenstellen. Zunächst sollten wir als Forscher positive Stereotypen und Essentials beschreiben und", "score": 64.0}1099{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_246.wav", "doc_id": "oYCKgTzTDy.seg_246", "src_text": "We found that Encoder-Decoder or Encoder-PTR can be improved by training in a mixture of various languages.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "und fanden heraus, dass Encoder-Decoder oder Encoder-PDR verbessert werden kann, indem sie in einer Mischung verschiedener Sprachen trainiert werden.", "score": 82.0}1100{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_640.wav", "doc_id": "FLkGnzVRew.seg_640", "src_text": "So why does this matter?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Warum ist das also wichtig?", "score": 100.0}1101{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_272.wav", "doc_id": "PIZEXUFLAR.seg_272", "src_text": "Here we present MultiInstruct, the first multi-modal instruction tuning benchmark dataset that consists of 62 diverse multi-modal tasks covering 10 broad categories.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hier stellen wir MultiInstrukt vor, das erste MultiModal Instruction Tuning Benchmark-Datensatz, der aus 62 verschiedenen MultiModalen Tasks besteht, die zehn Borencategorien abdecken.", "score": 21.0}1102{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_759.wav", "doc_id": "XejEJmgUmE.seg_759", "src_text": "And there we see that the MPP judgments either increase or decrease significantly when you add either acceptable prefixes or unacceptable prefixes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Dort sehen wir, dass die MP3-Urteile sich entweder erheblich oder unerheblich verändern, wenn man entweder akzeptable oder inakzeptable Präfixe hinzufügt.", "score": 55.0}1103{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_352.wav", "doc_id": "gGbuDbHhyc.seg_352", "src_text": "We can't stop on this problem setting, but this implies that additional manual annotations are required in weakly supervised learning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zweifel an dieser Problemstellung, da dies impliziert, dass zusätzliche manuelle Anmerkungen bei der Erstellung von Wochenplanen erforderlich sind,", "score": 66.0}1104{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_654.wav", "doc_id": "FLkGnzVRew.seg_654", "src_text": "We transfer from two different tasks: topic independent dissonance stance classification, a task that determines if two debate statements from different people are in agreement or in disagreement, irrespective of topic, called debate here, and on binary classification of expansion and comparison classes of PDTB since these two are closely related to the conception of consonance and dissonance and we call them CE here.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wir übertragen aus zwei verschiedenen Themen, unabhängige Themenklassifikation, ob zwei Erklärungen von verschiedenen Personen in Übereinstimmung oder im Widerspruch zum Thema stehen. Die Debatte hier und über die binäre Klassifizierung von Expansion und Vergleichsklassen von Pente, da diese eng mit der Konzeption von Konsonanten und Dissonanzen zusammenhängen, und wir nennen sie hier.", "score": 35.0}1105{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_54.wav", "doc_id": "TVCREhgqUP.seg_54", "src_text": "And \"Mary knew that the girl slept.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Mary wusste, dass das Mädchen geschlafen", "score": 70.0}1106{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_147.wav", "doc_id": "wLqFAuDnKa.seg_147", "src_text": "The dev data is much more curated, and with higher quality than the training data, that it's more noisy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die dev-Daten sind viel sorgfältiger bearbeitet und haben eine höhere Qualität als die Trainingsdaten, die lauter sind, und die Ergebnisse zeigen eine", "score": 90.0}1107{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_317.wav", "doc_id": "dJGfOSFgZO.seg_317", "src_text": "However, we believe there is a more precise and reliable strategy for dimensional dialogue evaluation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir glauben jedoch, dass es eine präzisere und zuverlässigere Strategie für die dimensionale Dialogbewertung gibt.", "score": 100.0}1108{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_320.wav", "doc_id": "dJGfOSFgZO.seg_320", "src_text": "We developed this method to comprehensively cover chat model behaviors that have been suggested to affect chat quality in recent literature.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir entwickelten diese Methode, um umfassend Chat-Modellverhaltens zu decken, die vorgeschlagen wurden, um Chat-Qualität und jüngste Literatur zu beeinflussen.", "score": 40.0}1109{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_363.wav", "doc_id": "gGbuDbHhyc.seg_363", "src_text": "Our second finding is that increasing the number of clean validation samples will help WSL approaches to achieve better performance, as shown in the figure on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Unsere zweite Erkenntnis ist, dass eine Erhöhung der Anzahl der Clean-Validation-Samples den WSL-Ansätzen helfen wird, bessere Leistungen zu erzielen, wie in der Abbildung auf der linken Seite gezeigt.", "score": 90.0}1110{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_854.wav", "doc_id": "GvEBWkLmuI.seg_854", "src_text": "Now for some results.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Manchmal haben wir Erfolge,", "score": 40.0}1111{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_383.wav", "doc_id": "WBLMIsdIrq.seg_383", "src_text": "Hello, my name is Kayo Yin and I will be presenting our work titled \"When Does Translation Require Context?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo, mein Name ist Kai-Ohin und ich werde unsere Arbeit präsentieren, die sich auf die Erforschung von Daten", "score": 40.0}1112{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_253.wav", "doc_id": "oYCKgTzTDy.seg_253", "src_text": "We found that, by comparing the green and orange line, we found the Zero-shot setting, the Cross-lingual transfer performance gap is significant, and then comparing the blue and orange lines, we found that with the Few-shot setting the transfer gap is shortened rapidly.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir fanden heraus, dass durch Vergleich der grünen und orangefarbenen Linie wir fanden, dass für die Einstellung ohne Schüsse die Leistungsdifferenz bei der Übersetzung zwischen Sprachen signifikant ist, und durch Vergleich der blauen und orangefarbenen Linie. Wir fanden heraus, dass mit wenigen Schritten die Transferlücke schnell geschlossen wird.", "score": 65.0}1113{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_211.wav", "doc_id": "SLpqvupgvW.seg_211", "src_text": "If the language model has access only to entity names, then the accuracy is only 60%, so there's a lot of room for improvement.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wenn das Sprachmodell nur Zugriff auf Entitätennamen hat, ist die Genauigkeit nur 60%, also gibt es viel Raum für Verbesserungen.", "score": 100.0}1114{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_542.wav", "doc_id": "dvGkKzmIaN.seg_542", "src_text": "As shown in the figures, it's hard to distinguish between, the backdoor embeddings and normal embeddings.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wie die Zahlen zeigen, ist es schwierig, zwischen den Backdoor-Embeddings und den normalen Embeddings zu unterscheiden.", "score": 100.0}1115{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_336.wav", "doc_id": "dJGfOSFgZO.seg_336", "src_text": "With the rapid pace of improvement in the field, many of these error rates could see a decrease in new models released since our evaluation was conducted.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Mit dem schnellen Tempo der Verbesserung in diesem Bereich könnten viele dieser Fehler in neuen Modellen, die veröffentlicht werden, abnehmen.", "score": 95.0}1116{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_521.wav", "doc_id": "dvGkKzmIaN.seg_521", "src_text": "Watermark injection and copyright verification.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wasserzeichen-Einspritzung und Urheberrechtsanwendung. Bevor wir diese", "score": 40.0}1117{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_324.wav", "doc_id": "dJGfOSFgZO.seg_324", "src_text": "For comparison, we also evaluated these conversations using three existing methods: Likert ratings on the turn-level, Likert ratings on the dialogue-level, and dialogue-level pairwise comparisons.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Vergleich haben wir diese Gespräche auch mit drei bestehenden Methoden bewertet. Lickert-Bewertungen auf der Drehstufe, Lickert-Bewertungen auf der Dialogstufe und Dialogstufe-paareisen Vergleiche.", "score": 60.0}1118{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_433.wav", "doc_id": "hgIDlKNiFM.seg_433", "src_text": "Then we will present the main contribution of our article.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "stellen wir die wichtigsten Beiträge unserer Arbeit vor:", "score": 85.0}1119{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_10.wav", "doc_id": "aQpIWggfCo.seg_10", "src_text": "In this paper, we first evaluate and improve the constrained language planning ability of large language models.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In dieser Arbeit bewerten wir zunächst und verbessern die konstruktionsbezogene Sprachplanungsfähigkeit von großen Sprachmodellen.", "score": 90.0}1120{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_203.wav", "doc_id": "SLpqvupgvW.seg_203", "src_text": "Here are some examples from our dataset.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hier sind einige Beispiele aus unserem Datensatz.", "score": 100.0}1121{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_814.wav", "doc_id": "WTTtiRKFZI.seg_814", "src_text": "Right?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "ist der", "score": 5.0}1122{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_457.wav", "doc_id": "hgIDlKNiFM.seg_457", "src_text": "However, our experiment on control pre-training using the weight and tokenization of CamemBERT trained on the four GB subset of NACHOS showed comparable results to those obtained with DrBERT 4 GB from-scratch.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Weiterbildung, bei denen wir das Gewicht und den Tokenizer von Pamper Bert verwenden, und trainieren, um auf dem 4-GB-Untersatz von Natures zu trainieren, zeigen vergleichbare Ergebnisse, die wir mit dem Doktor Bert von Scratch erzielt haben.", "score": 30.0}1123{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_50.wav", "doc_id": "TVCREhgqUP.seg_50", "src_text": "Compositional generalization can be understood as the ability of a learner to handle deeper recursion and unseen compositions of phrases that have been seen individually during training.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Die kompositorische Generalisierung kann als die Fähigkeit eines Lernenden verstanden werden, tiefere Wiederholungen und unsichtbare Kompositionen von Aussagen zu handhaben, die während des Trainings individuell gesehen wurden.", "score": 85.0}1124{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_465.wav", "doc_id": "SUkmfOTvGi.seg_465", "src_text": "Let's get started.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Sie uns anfangen.", "score": 90.0}1125{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_802.wav", "doc_id": "WTTtiRKFZI.seg_802", "src_text": "When you swap these two constituents, the sum of these two dependencies becomes 6.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "und sie austauschen, wird die Summe dieser beiden Abhängigkeiten sechs", "score": 60.0}1126{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_742.wav", "doc_id": "XejEJmgUmE.seg_742", "src_text": "So what we do is that to simulate these longer sequences, we revisit the data sets themselves and then we recreate sentences by choosing acceptable or unacceptable sentences from those datasets.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "den wir verwenden, um diese längeren Sequenzen zu simulieren: Wir überprüfen die Datensätze selbst und erstellen dann Sätze aus akzeptablen oder unakzeptablen Sätzen.", "score": 90.0}1127{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_877.wav", "doc_id": "GvEBWkLmuI.seg_877", "src_text": "We just really can't make any assumptions or really study that further, without more transparency.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "können wir wirklich keine Annahmen treffen oder sie weiter studieren, ohne mehr Transparenz.", "score": 70.0}1128{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_623.wav", "doc_id": "oeooqChmKK.seg_623", "src_text": "Without task-specific training on KITMUS, both models do not perform well.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Modelle schnitten beim spezifischen Training auf dem Kidus nicht gut ab,", "score": 60.0}1129{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_554.wav", "doc_id": "rISrKoXQCx.seg_554", "src_text": "To this end, we propose to investigate the political bias propagation pipeline from pretraining data to language models to downstream tasks, specifically by asking the following questions: First, how do we evaluate the political leaning of language models and what role does pretraining data might have on such political biases?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zu diesem Zweck schlagen wir vor, die politische Bias-Verbreitungspipeline von der Vorverarbeitung von Daten zu Sprachmodellen zu Downstream-Aufgaben zu untersuchen, insbesondere, indem wir die folgenden Fragen stellen: Zunächst, wie bewerten wir die politische Ausrichtung von Sprachmodellen und welche Rolle könnte die Vorverarbeitung von Daten auf solche politischen Vorurteile haben?", "score": 85.0}1130{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_787.wav", "doc_id": "WTTtiRKFZI.seg_787", "src_text": "The argument is based on the principle of dependency length minimization that I will explain on the basis of these examples.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Argument basiert auf dem Prinzip der Abhängigkeit der Minimierung, das ich auf der Grundlage dieser Beispiele erkläre. So", "score": 70.0}1131{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_480.wav", "doc_id": "SUkmfOTvGi.seg_480", "src_text": "The second ingredient is the model size.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Der zweite Bestandteil ist die Modellgröße.", "score": 100.0}1132{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_124.wav", "doc_id": "wLqFAuDnKa.seg_124", "src_text": "PaLM is a 540 billion-parameter large language model presented last year in 2022.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "PP-1M ist ein 540 Milliarden Parameter großes Sprachmodell, das im Jahr 2022 vorgestellt wurde.", "score": 40.0}1133{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_760.wav", "doc_id": "XejEJmgUmE.seg_760", "src_text": "But when we match the structure, that is when we choose the sentences from the same phenomena in BLiMP or SyntaxGym, we see a massive increase or a massive decrease of the MPP judgement for the model, depending on whether the chosen prefix is acceptable or unacceptable.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Aber wenn wir die Struktur übereinstimmen lassen, dann ist das der Zeitpunkt, an dem wir die Sätze aus dem gleichen Phänomen im Blimp-Syntax auswählen. Wir sehen eine massive Zunahme oder Abnahme des MPP-Urteils für das Modell, je nachdem, ob das gewählte Präfix akzeptabel oder nicht akzeptabel ist.", "score": 65.0}1134{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_780.wav", "doc_id": "WTTtiRKFZI.seg_780", "src_text": "The conjunction headed approach assumed in Prague dependency treebanks, where coordinate structures are headed by the conjunction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Prag-Ansatz und der Konjunktions-Ansatz, die Koordinatenstrukturen, die von der Konjunktion abhängen. Daher", "score": 50.0}1135{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_137.wav", "doc_id": "wLqFAuDnKa.seg_137", "src_text": "So, it's important to select a good prompting strategy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "erreichen, weshalb es wichtig ist, eine gute Strategie zu wählen.", "score": 85.0}1136{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_710.wav", "doc_id": "oaOHnMCwad.seg_710", "src_text": "So for the GPT 4 social acceptability analysis, we find that it's most aligned to confucian and English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "und für die GPD-Analyse finden wir heraus, dass die meisten englischsprachigen Länder am besten geeignet sind. Wir finden", "score": 59.0}1137{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_639.wav", "doc_id": "FLkGnzVRew.seg_639", "src_text": "While dissonance is a very common phenomenon we experienced in daily decision making, they are really rare to find expressed in language among other kinds of discourse relations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Während Dissens ein sehr häufiges Phänomen ist, das wir in der täglichen Entscheidungsfindung erleben, sind sie in der Sprache in anderen Arten von Diskursbeziehungen wirklich selten ausgedrückt.", "score": 65.0}1138{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_92.wav", "doc_id": "uZBWfYjYnf.seg_92", "src_text": "Hi, I'm Sara Papi from the University of Trento and Foundazione Bruno Kessler and I will briefly introduce the \"Attention as a Guide for Simultaneous Speech Translation\" paper, that is a joint work with Matteo Negri and Marco Turchi.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo, ich bin Sera Papí von der Universität von Trento und der Bruno Kessler Foundation, und ich werde kurz die Aufmerksamkeit als Leitfaden für ein Simultansprachübersetzungs-Papier vorstellen, das eine gemeinsame Arbeit mit Matteo Negri und Marco Turci ist.", "score": 70.0}1139{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_2.wav", "doc_id": "aQpIWggfCo.seg_2", "src_text": "In everyday life, humans often plan their actions by following step-by-step instructions in the form of goal-oriented scripts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Im Alltag planen Menschen oft ihre Handlungen, indem sie Schritt für Schritt Anweisungen in Form von orientierten Skripten befolgen.", "score": 85.0}1140{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/aQpIWggfCo.seg_35.wav", "doc_id": "aQpIWggfCo.seg_35", "src_text": "In total, we generate 55,000 specific goals with scripts.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Insgesamt generieren wir 50.000 spezifische Ziele mit Skripten,", "score": 85.0}1141{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_652.wav", "doc_id": "FLkGnzVRew.seg_652", "src_text": "To alleviate this, we experiment over combinations of transfer learning and active learning to annotate such that more dissonant samples can be collected over lesser annotation runs, lowering the overall annotation costs while improving dissonance detection.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Um dies zu erleichtern, experimentieren wir mit Kombinationen von Transferlernen und aktiven Lernvorgängen, um solche Anmerkungen zu annotieren, sodass mehr Dissensbeispiele über weniger Anmerkungsrunden gesammelt werden können, wodurch die Gesamtkosten der Anmerkung verringert und die Dissensdetektion verbessert werden.", "score": 95.0}1142{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_825.wav", "doc_id": "WTTtiRKFZI.seg_825", "src_text": "So see the paper for the full arguments.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Sehen Sie sich das Papier für die vollständige", "score": 80.0}1143{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/rISrKoXQCx.seg_554.wav", "doc_id": "rISrKoXQCx.seg_554", "src_text": "To this end, we propose to investigate the political bias propagation pipeline from pretraining data to language models to downstream tasks, specifically by asking the following questions: First, how do we evaluate the political leaning of language models and what role does pretraining data might have on such political biases?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zu diesem Zweck schlagen wir vor, die politische Verbreitungspipeline zu untersuchen, indem Daten vorbereitet werden, Sprachmodelle erstellt werden, Downstream-Tasks durchgeführt werden, insbesondere durch die Stellung der folgenden Fragen. Erstens: Wie können wir die politische Tendenz von Sprachmodellen bewerten, und welche Rolle könnte Perusini-Daten auf solche politischen Vorurteile haben? Zweitens,", "score": 64.0}1144{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_122.wav", "doc_id": "wLqFAuDnKa.seg_122", "src_text": "Hello everyone, my name is David Vilar, and I will be giving a short review of the paper \"Prompting PaLM for Translation: Assessing Strategies and Performance.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hallo, mein Name ist Vilar und ich werde eine kurze Zusammenfassung des Papiers zu Übersetzung, Bewertung, Strategien und Leistung geben.", "score": 60.0}1145{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_162.wav", "doc_id": "SLpqvupgvW.seg_162", "src_text": "Our goal is to understand users’ language when they want to make a choice.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Unser Ziel ist es, die Sprache der Benutzer zu verstehen, wenn sie eine Wahl treffen", "score": 100.0}1146{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_665.wav", "doc_id": "FLkGnzVRew.seg_665", "src_text": "On further rounds of AL with two best strategies, we improve dissonance classification AUC to 0.75, which is the best performance that we have on the task so far.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Bei den nächsten Runden mit zwei besseren Strategien verbessern wir die Diskriminanzklassifizierung von AUC zu einem Punkt von sieben, was die beste Leistung ist, die wir bisher hatten.", "score": 55.0}1147{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_406.wav", "doc_id": "WBLMIsdIrq.seg_406", "src_text": "And this allows us to find, for example, dual pronouns in Arabic that have relatively high P-CXMI.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "sich zum Beispiel anhand von Pronomen in Arabisch bestimmen, die einen hohen Index haben,", "score": 42.0}1148{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_378.wav", "doc_id": "gGbuDbHhyc.seg_378", "src_text": "Third, continuous fine-tuning is a simple yet strong baseline that should be considered in future work in WSL.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Drittens ist die kontinuierliche Feinabstimmung eine einfache, aber starke Grundlage, die in zukünftigen Arbeiten in WSL berücksichtigt werden sollte.", "score": 100.0}1149{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_263.wav", "doc_id": "PIZEXUFLAR.seg_263", "src_text": "Hello everyone, my name is Ying and my colleague Zhiyang and I will be presenting our research on MultiInstruct improving Multi-Modal Zero-Shot Learning via Instruction Tuning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Hallo, ich heiße Yen und ich werde zusammen mit meinem Kollegen Jiajun unsere Forschung über Multi-Instruction präsentieren, die multimodale neuronale Lernfähigkeit durch Anpassung der Anweisungen verbessert.", "score": 60.0}1150{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_525.wav", "doc_id": "dvGkKzmIaN.seg_525", "src_text": "In watermark injection, we first define a target embedding.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wasserzeicheninjektion definieren wir zunächst eine Zielverankerung.", "score": 100.0}1151{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_178.wav", "doc_id": "SLpqvupgvW.seg_178", "src_text": "And with that, Bob sets the dialogue context.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Und mit dieser Aussage setzt Bob den Dialogkontext.", "score": 100.0}1152{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/uZBWfYjYnf.seg_100.wav", "doc_id": "uZBWfYjYnf.seg_100", "src_text": "So what is our solution?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Was ist unsere Lösung?", "score": 95.0}1153{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_856.wav", "doc_id": "GvEBWkLmuI.seg_856", "src_text": "However, when we actually look at the distribution of the words and lexicon, we find very different things.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wenn wir jedoch die Verteilung der Wörter im Lexikon betrachten, finden wir sehr unterschiedliche Dinge, sodass", "score": 95.0}1154{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_383.wav", "doc_id": "WBLMIsdIrq.seg_383", "src_text": "Hello, my name is Kayo Yin and I will be presenting our work titled \"When Does Translation Require Context?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Hallo, mein Name ist Kayen und ich werde unsere Arbeit mit dem Titel „Windows-Übersetzungskontext: Eine mehrsprachige Untersuchung“", "score": 50.0}1155{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_358.wav", "doc_id": "gGbuDbHhyc.seg_358", "src_text": "We addressed these research questions in our work and our findings are as follows.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir behandeln diese Forschungsfragen in unserer Arbeit und unsere Ergebnisse sind wie folgt. Zunächst stellen", "score": 95.0}1156{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/PIZEXUFLAR.seg_279.wav", "doc_id": "PIZEXUFLAR.seg_279", "src_text": "Ok, now I'm going to talk about multi-modal instruction tuning.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Okay, nun werde ich über die Abstimmung der Multimode-Anweisung sprechen.", "score": 55.0}1157{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_775.wav", "doc_id": "WTTtiRKFZI.seg_775", "src_text": "A similar approach is assumed in Igor Mel'čuk's meaning text theory, where again, the whole coordinate structure is headed by the first conjuct.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "wird durch die gesamte korrelierte Struktur bestimmt, so dass diese beiden", "score": 40.0}1158{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_539.wav", "doc_id": "dvGkKzmIaN.seg_539", "src_text": "The results on four data sets show that our embedding marker can have great detection performance while keep great utility for downstream tasks.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Die Ergebnisse auf vier Datensätzen zeigen, dass unser eingebetteter Marker eine ausgezeichnete Erkennungsleistung aufweist und gleichzeitig eine ausgezeichnete Nützlichkeit für Downstream-Aufgaben hat.", "score": 100.0}1159{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_868.wav", "doc_id": "GvEBWkLmuI.seg_868", "src_text": "And finally, for black women, we see that some of the top words are things like \"strong\" and \"resilient\".", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "schließlich sehen wir bei schwarzen Frauen, dass einige der obersten Wörter Dinge sind, die stark und widerstandsfähig sind.", "score": 85.0}1160{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_174.wav", "doc_id": "SLpqvupgvW.seg_174", "src_text": "Our data set covers three different domains: music, books, and recipes.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Unser Datensatz umfasst drei verschiedene Bereiche: Musik, Bücher und Rezensionen.", "score": 80.0}1161{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_710.wav", "doc_id": "oaOHnMCwad.seg_710", "src_text": "So for the GPT 4 social acceptability analysis, we find that it's most aligned to confucian and English speaking countries.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "So finden wir heraus, dass die Datensätze und Modelle für die GPD4-Social-Acceptability-Analyse am meisten mit Konfuzianismus und englischsprachigen Ländern ausgerichtet sind.", "score": 66.0}1162{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_123.wav", "doc_id": "wLqFAuDnKa.seg_123", "src_text": "This is joint work with my colleagues from Google Translate.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dies ist eine gemeinsame Arbeit mit meinen Kollegen von Google Translate.", "score": 100.0}1163{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_320.wav", "doc_id": "dJGfOSFgZO.seg_320", "src_text": "We developed this method to comprehensively cover chat model behaviors that have been suggested to affect chat quality in recent literature.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir entwickeln diese Methode, um Verhaltensweisen in Chat-Modellen abzubilden, die sich auf die Chat-Qualität", "score": 67.0}1164{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/GvEBWkLmuI.seg_853.wav", "doc_id": "GvEBWkLmuI.seg_853", "src_text": "So for instance, for the personas of black women, we would do Fightin’ Words and compare the log-odds ratios against both white personas and man personas because those are the two corresponding unmarked groups.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Zum Beispiel für die Personas von schwarzen Frauen würden wir Fighting Words verwenden und die Logits-Raten gegenüber sowohl weißen Personas als auch männlichen Personas vergleichen, weil es sich um die zwei korrespondierenden unmarkierten Gruppen handelt.", "score": 80.0}1165{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_369.wav", "doc_id": "gGbuDbHhyc.seg_369", "src_text": "As we can see from the figures, the vanilla model, termed FTw, initially underperforms more complicated WSL methods, like COSINE.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Wie aus den Zahlen ersichtlich, schneidet das Wallina-Modell mit der Bezeichnung FTW anfangs schlechter ab als komplexere WSL-Methoden wie Kosinus.", "score": 90.0}1166{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_770.wav", "doc_id": "XejEJmgUmE.seg_770", "src_text": "Thank you for listening.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Vielen Dank für Ihre Aufmerksamkeit.", "score": 100.0}1167{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_815.wav", "doc_id": "WTTtiRKFZI.seg_815", "src_text": "So the governor is on the left in this example \"I saw Bart and Lisa\" so is the governor is on the left.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "diesem Beispiel der Gouverneur links ist, aber nicht in dem", "score": 30.0}1168{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_457.wav", "doc_id": "hgIDlKNiFM.seg_457", "src_text": "However, our experiment on control pre-training using the weight and tokenization of CamemBERT trained on the four GB subset of NACHOS showed comparable results to those obtained with DrBERT 4 GB from-scratch.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "meisten Aufgaben zu führen. Unsere Experimente mit kontinuierlicher Gewichtsabtastung mit dem Gewichts- und Token-System von Pommert zeigen vergleichbare Ergebnisse mit denen von Dr. Pommert.", "score": 40.0}1169{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_149.wav", "doc_id": "wLqFAuDnKa.seg_149", "src_text": "Nevertheless, specialized state-of-the-art systems have a substantial advantage over the PaLM translations.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "haben spezialisierte Systeme einen erheblichen Vorteil gegenüber den", "score": 1.0}1170{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_439.wav", "doc_id": "hgIDlKNiFM.seg_439", "src_text": "Since then, this model has been adapted to many other languages, like in French with CamemBERT, and also in domains like biomedical with PubMedBERT and BioBERT and on clinical with ClinicalBERT, but mostly in English.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Seitdem wurde dieses Modell an viele andere Sprachen angepasst, z. B. an Französisch mit Camber und andere Domänen wie Biomedizin mit Pamet und Biot sowie klinisch mit klinischen Begriffen, aber größtenteils auf Englisch. Spezialisierte Modelle für", "score": 35.0}1171{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_618.wav", "doc_id": "oeooqChmKK.seg_618", "src_text": "In the Background-Pretrain setting, we assume that the background knowledge \"Politicians seek elected seats in government\" is contained in the pretrained parameters and in inference-time context we provide the entity-specific knowledge \"Chichester is a politician.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Im Hintergrundwissen nehmen wir an, dass die Hintergrundkenntnisse von Politikern, die Sitze in der Regierung anstreben, in den Hintergrundparametern enthalten sind. Im Kontext der Gegenwart liefern wir das spezifische Wissen, dass Chester ein Politiker", "score": 70.0}1172{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/gGbuDbHhyc.seg_342.wav", "doc_id": "gGbuDbHhyc.seg_342", "src_text": "In this video, I would like to present our recent work \"Weaker Than You Think: A Critical Look at Weakly Supervised Learning.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "In diesem Video möchte ich unsere jüngste Arbeit präsentieren, eine kritische Betrachtung der wöchentlichen Nachrichten.", "score": 20.0}1173{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/hgIDlKNiFM.seg_449.wav", "doc_id": "hgIDlKNiFM.seg_449", "src_text": "Another also based on CamemBERT, but trained this time on the 4 GB of clinical notes and finally, one based on English biomedical model PubMedBERT, and trained on 4 GB of set of NACHOS.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "dem Gewicht von Camembert, aber trainiert diesmal auf vier Gigabyte von Klicks. Insgesamt haben wir sieben Modelle. Um alle sieben Modelle zu bewerten, haben wir öffentliche und private Aufgaben wie Name-Recognition, Klassifizierung, Part-of-Speech-Tagging und Fragen beantworten. Dieses Modell entspricht sechs Zeilen des Modells,", "score": 30.0}1174{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_702.wav", "doc_id": "oaOHnMCwad.seg_702", "src_text": "Afterwards to stay engaged in the study, they can compare their responses to an AI and others.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Danach können sie die Antworten von A und anderen vergleichen.", "score": 38.0}1175{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_744.wav", "doc_id": "XejEJmgUmE.seg_744", "src_text": "And what we do is that to recreate like longer sequences and which are acceptable and which has the same matching of the grammatical structure.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "In unserem Fall erzeugen wir längerer Sequenzen, die akzeptabel sind und die gleiche grammatische Struktur haben,", "score": 92.0}1176{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SUkmfOTvGi.seg_480.wav", "doc_id": "SUkmfOTvGi.seg_480", "src_text": "The second ingredient is the model size.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Der zweite Bestandteil ist die Modellgröße.", "score": 98.0}1177{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_650.wav", "doc_id": "FLkGnzVRew.seg_650", "src_text": "To no surprise, the classifier performed not much better than chance.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Es war kein Wunder, dass der Klassifikator nicht viel besser als zufällig leistete.", "score": 97.0}1178{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WTTtiRKFZI.seg_780.wav", "doc_id": "WTTtiRKFZI.seg_780", "src_text": "The conjunction headed approach assumed in Prague dependency treebanks, where coordinate structures are headed by the conjunction.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "den Konjunktionskopf-Ansatz in pragmatischen Abhängigkeitstrees, wo koordinierte Strukturen vom Konjunkt angeführt werden.", "score": 44.0}1179{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/FLkGnzVRew.seg_641.wav", "doc_id": "FLkGnzVRew.seg_641", "src_text": "Studying cognitive dissonance can help us understand the effects of disagreement among people, track trends and belief values, and attitude changes in population.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Die Untersuchung kognitiver Differenzen kann uns helfen, die Auswirkungen von Meinungsverschiedenheiten unter Menschen zu verstehen, Trends in Überzeugungen, Werten und Einstellungen in der Bevölkerung zu verfolgen und die Auswirkungen von sozialen und kulturellen Veränderungen auf die kognitive Landschaft zu verstehen.", "score": 82.0}1180{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_249.wav", "doc_id": "oYCKgTzTDy.seg_249", "src_text": "We also compare the cross-language performance gap.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir vergleichen auch die Cross-Lang Performance Gap.", "score": 85.0}1181{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_322.wav", "doc_id": "dJGfOSFgZO.seg_322", "src_text": "For example, ABC-Eval measures the number of turns in which a chat model ignores its partner or says something irrelevant, contradicts itself or its partner, hallucinates incorrect facts or violates common sense knowledge, and when the model succeeds or fails to show empathy.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zum Beispiel misst A.B.C. E.V.A.L. die Anzahl der Drehungen, in denen ein Chat-Modell seinen Partner ignoriert oder etwas Irrelevantes sagt. Widerspricht sich selbst oder seiner Partner, halluziniert falsche Tatsachen oder verletzt das Wissen des Menschen, und wenn das Modell die Empathie zeigt oder nicht zeigt, ist es erfolgreich oder scheitert.", "score": 68.0}1182{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_402.wav", "doc_id": "WBLMIsdIrq.seg_402", "src_text": "Now we analyze words with high P-CXMI to look for patterns between these words.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "sind. Jetzt analysieren wir die Wörter mit hoher Häufigkeit, um die Paare zwischen diesen Wörtern zu finden.", "score": 60.0}1183{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oYCKgTzTDy.seg_262.wav", "doc_id": "oYCKgTzTDy.seg_262", "src_text": "Thanks for listening.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "fürs Zuhören.", "score": 90.0}1184{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_541.wav", "doc_id": "dvGkKzmIaN.seg_541", "src_text": "The legend of the figures means the number of triggers in each sentence.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Die Legende der Figuren bedeutet die Anzahl der Auslöser in jedem Satz.", "score": 90.0}1185{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dJGfOSFgZO.seg_334.wav", "doc_id": "dJGfOSFgZO.seg_334", "src_text": "For example, the bots we tested have common sense violations in around 20% of their responses.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Zum Beispiel haben die Roboter, die wir getestet haben, in etwa 20% ihrer Antworten Verstöße gegen den Grundsatz der Menschenwürde gezeigt.", "score": 20.0}1186{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_77.wav", "doc_id": "TVCREhgqUP.seg_77", "src_text": "Then we jump to the next multiset token, to determine the second token in the output.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Dann springen wir zum nächsten Multiset-Token, um den zweiten Token im Output zu bestimmen.", "score": 94.0}1187{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_186.wav", "doc_id": "SLpqvupgvW.seg_186", "src_text": "Do you mean A or B?", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Meinen Sie A oder B?", "score": 100.0}1188{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_122.wav", "doc_id": "wLqFAuDnKa.seg_122", "src_text": "Hello everyone, my name is David Vilar, and I will be giving a short review of the paper \"Prompting PaLM for Translation: Assessing Strategies and Performance.\"", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Hallo Irland, mein Name ist Aidan Villar und ich werde Ihnen eine kurze Zusammenfassung des Papiers vorstellen, das sich mit der Leistungsfähigkeit der Übersetzung befasst, die Strategien und Leistung.", "score": 31.0}1189{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_598.wav", "doc_id": "oeooqChmKK.seg_598", "src_text": "We introduce a coreference resolution task, designed to probe for the ability to draw on knowledge available in different sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "Wir stellen eine Korreferenz-Resolution-Aufgabe vor, die dazu konzipiert ist, die Fähigkeit zu testen, auf Kenntnisse in verschiedenen Quellen zurückzugreifen:", "score": 89.0}1190{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/XejEJmgUmE.seg_745.wav", "doc_id": "XejEJmgUmE.seg_745", "src_text": "We extract grammatical sentences from Adjunct Island and then we add it as a prefix to both the acceptable query and the unacceptable query.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "long_KIT_primary", "tgt_text": "indem wir Grammatik-Sätze aus dem Adversarial-Tool extrahieren und sie als Präfix zu sowohl der akzeptablen als auch der nicht akzeptablen Frage", "score": 65.0}1191{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/SLpqvupgvW.seg_173.wav", "doc_id": "SLpqvupgvW.seg_173", "src_text": "We're not aware of a larger-scale public data set for the task, so we collect one using crowd annotation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Wir sind uns nicht bewusst, dass es ein öffentliches Datensatzsystem gibt, ein groß angelegtes Datensatzsystem für Aufgaben, also sammeln wir einen, indem wir die Minderheit benutzen;", "score": 43.0}1192{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/dvGkKzmIaN.seg_510.wav", "doc_id": "dvGkKzmIaN.seg_510", "src_text": "To protect the copyright of embedding as services, one of the solutions is to embed a watermark in the provider service and detect whether another service contain the watermark.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Um die Urheberrechte an eingebetteten Diensten zu schützen, ist eine der Lösungen, ein Wasserzeichen in den Dienst des Anbieters einzubetten und zu prüfen, ob ein anderer Dienst das Wasserzeichen enthält.", "score": 100.0}1193{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_146.wav", "doc_id": "wLqFAuDnKa.seg_146", "src_text": "In particular, we compare the selecting prompts from the training data for the WMT evaluations on the dev data.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "hoher Qualität auszuwählen, insbesondere die aus den Trainingsdaten der WMT-Evaluierungen oder den Testdaten.", "score": 30.0}1194{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_124.wav", "doc_id": "wLqFAuDnKa.seg_124", "src_text": "PaLM is a 540 billion-parameter large language model presented last year in 2022.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Parm ist ein fünfundvierzig Milliarden Parameter großes Sprachmodell, das im vergangenen Jahr vorgestellt wurde.", "score": 100.0}1195{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/TVCREhgqUP.seg_65.wav", "doc_id": "TVCREhgqUP.seg_65", "src_text": "Obtaining trees may also involve specialized grammar-induction procedures.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "von Bäumen kann auch spezielle Grammatikinduktionsprozesse beinhalten.", "score": 88.0}1196{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_153.wav", "doc_id": "wLqFAuDnKa.seg_153", "src_text": "So, in particular, the most common errors are omission errors.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_primary", "tgt_text": "Insbesondere sind die häufigsten Fehler Auslassungsfehler.", "score": 100.0}1197{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oaOHnMCwad.seg_701.wav", "doc_id": "oaOHnMCwad.seg_701", "src_text": "We host 2 tasks on lab in the wild, one of them being social acceptability, and the way this works is that participants will read a situation from the social chemistry dataset and, then they'll write how socially acceptable a situation is.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Wir führen zwei Tests durch, die soziale Akzeptanz zu ermitteln, und die Art und Weise, wie diese Tests durchgeführt werden, ist", "score": 40.0}1198{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/oeooqChmKK.seg_617.wav", "doc_id": "oeooqChmKK.seg_617", "src_text": "Here's an example of how we control the availability of facts in the true sources.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Hier ist ein Beispiel dafür, wie man die Verfügbarkeit von Fakten in echten Quellen steuern kann.", "score": 100.0}1199{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/WBLMIsdIrq.seg_391.wav", "doc_id": "WBLMIsdIrq.seg_391", "src_text": "However, evaluating how well models can translate cases like this is pretty hard.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_CUNI-NL_contrastive", "tgt_text": "Übersetzung ändert sich ebenfalls. Die Bewertung, wie gut Modelle solche Fälle übersetzen können,", "score": 100.0}1200{"audio_path": "data/iwslt25/IWSLT25INSTRUCT/segmented/wLqFAuDnKa.seg_154.wav", "doc_id": "wLqFAuDnKa.seg_154", "src_text": "So, it seems that PaLM chooses to produce a better-sounding translation, sometimes by dropping parts of the source sentence that are made in translation.", "src_text_system": "human", "src_lang": "en", "tgt_lang": "de", "domain": "acl", "tgt_system": "short_NLE_primary", "tgt_text": "Es scheint, dass Palme sich dafür entscheidet, eine bessere Übersetzung zu produzieren, indem sie manchmal Teile der verarbeiteten Satzzeilen aus der Übersetzung entfernt.", "score": 80.0}