bigscience/mt0-small
325.3k
1---2datasets:3- bigscience/xP34- mc45license: apache-2.06language:7- af8- am9- ar10- az11- be12- bg13- bn14- ca15- ceb16- co17- cs18- cy19- da20- de21- el22- en23- eo24- es25- et26- eu27- fa28- fi29- fil30- fr31- fy32- ga33- gd34- gl35- gu36- ha37- haw38- hi39- hmn40- ht41- hu42- hy43- ig44- is45- it46- iw47- ja48- jv49- ka50- kk51- km52- kn53- ko54- ku55- ky56- la57- lb58- lo59- lt60- lv61- mg62- mi63- mk64- ml65- mn66- mr67- ms68- mt69- my70- ne71- nl72- no73- ny74- pa75- pl76- ps77- pt78- ro79- ru80- sd81- si82- sk83- sl84- sm85- sn86- so87- sq88- sr89- st90- su91- sv92- sw93- ta94- te95- tg96- th97- tr98- uk99- und100- ur101- uz102- vi103- xh104- yi105- yo106- zh107- zu108pipeline_tag: text2text-generation109widget:110- text: "一个传奇的开端,一个不灭的神话,这不仅仅是一部电影,而是作为一个走进新时代的标签,永远彪炳史册。Would you rate the previous review as positive, neutral or negative?"111 example_title: "zh-en sentiment"112- text: "一个传奇的开端,一个不灭的神话,这不仅仅是一部电影,而是作为一个走进新时代的标签,永远彪炳史册。你认为这句话的立场是赞扬、中立还是批评?"113 example_title: "zh-zh sentiment"114- text: "Suggest at least five related search terms to \"Mạng neural nhân tạo\"."115 example_title: "vi-en query"116- text: "Proposez au moins cinq mots clés concernant «Réseau de neurones artificiels»."117 example_title: "fr-fr query"118- text: "Explain in a sentence in Telugu what is backpropagation in neural networks."119 example_title: "te-en qa"120- text: "Why is the sky blue?"121 example_title: "en-en qa"122- text: "Write a fairy tale about a troll saving a princess from a dangerous dragon. The fairy tale is a masterpiece that has achieved praise worldwide and its moral is \"Heroes Come in All Shapes and Sizes\". Story (in Spanish):"123 example_title: "es-en fable"124- text: "Write a fable about wood elves living in a forest that is suddenly invaded by ogres. The fable is a masterpiece that has achieved praise worldwide and its moral is \"Violence is the last refuge of the incompetent\". Fable (in Hindi):"125 example_title: "hi-en fable"126model-index:127- name: mt0-small128 results:129 - task:130 type: Coreference resolution131 dataset:132 type: winogrande133 name: Winogrande XL (xl)134 config: xl135 split: validation136 revision: a80f460359d1e9a67c006011c94de42a8759430c137 metrics:138 - type: Accuracy139 value: 50.51140 - task:141 type: Coreference resolution142 dataset:143 type: Muennighoff/xwinograd144 name: XWinograd (en)145 config: en146 split: test147 revision: 9dd5ea5505fad86b7bedad667955577815300cee148 metrics:149 - type: Accuracy150 value: 51.31151 - task:152 type: Coreference resolution153 dataset:154 type: Muennighoff/xwinograd155 name: XWinograd (fr)156 config: fr157 split: test158 revision: 9dd5ea5505fad86b7bedad667955577815300cee159 metrics:160 - type: Accuracy161 value: 54.22162 - task:163 type: Coreference resolution164 dataset:165 type: Muennighoff/xwinograd166 name: XWinograd (jp)167 config: jp168 split: test169 revision: 9dd5ea5505fad86b7bedad667955577815300cee170 metrics:171 - type: Accuracy172 value: 52.45173 - task:174 type: Coreference resolution175 dataset:176 type: Muennighoff/xwinograd177 name: XWinograd (pt)178 config: pt179 split: test180 revision: 9dd5ea5505fad86b7bedad667955577815300cee181 metrics:182 - type: Accuracy183 value: 51.71184 - task:185 type: Coreference resolution186 dataset:187 type: Muennighoff/xwinograd188 name: XWinograd (ru)189 config: ru190 split: test191 revision: 9dd5ea5505fad86b7bedad667955577815300cee192 metrics:193 - type: Accuracy194 value: 54.29195 - task:196 type: Coreference resolution197 dataset:198 type: Muennighoff/xwinograd199 name: XWinograd (zh)200 config: zh201 split: test202 revision: 9dd5ea5505fad86b7bedad667955577815300cee203 metrics:204 - type: Accuracy205 value: 54.17206 - task:207 type: Natural language inference208 dataset:209 type: anli210 name: ANLI (r1)211 config: r1212 split: validation213 revision: 9dbd830a06fea8b1c49d6e5ef2004a08d9f45094214 metrics:215 - type: Accuracy216 value: 34.7217 - task:218 type: Natural language inference219 dataset:220 type: anli221 name: ANLI (r2)222 config: r2223 split: validation224 revision: 9dbd830a06fea8b1c49d6e5ef2004a08d9f45094225 metrics:226 - type: Accuracy227 value: 34.0228 - task:229 type: Natural language inference230 dataset:231 type: anli232 name: ANLI (r3)233 config: r3234 split: validation235 revision: 9dbd830a06fea8b1c49d6e5ef2004a08d9f45094236 metrics:237 - type: Accuracy238 value: 33.83239 - task:240 type: Natural language inference241 dataset:242 type: super_glue243 name: SuperGLUE (cb)244 config: cb245 split: validation246 revision: 9e12063561e7e6c79099feb6d5a493142584e9e2247 metrics:248 - type: Accuracy249 value: 50.0250 - task:251 type: Natural language inference252 dataset:253 type: super_glue254 name: SuperGLUE (rte)255 config: rte256 split: validation257 revision: 9e12063561e7e6c79099feb6d5a493142584e9e2258 metrics:259 - type: Accuracy260 value: 61.01261 - task:262 type: Natural language inference263 dataset:264 type: xnli265 name: XNLI (ar)266 config: ar267 split: validation268 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16269 metrics:270 - type: Accuracy271 value: 37.43272 - task:273 type: Natural language inference274 dataset:275 type: xnli276 name: XNLI (bg)277 config: bg278 split: validation279 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16280 metrics:281 - type: Accuracy282 value: 37.55283 - task:284 type: Natural language inference285 dataset:286 type: xnli287 name: XNLI (de)288 config: de289 split: validation290 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16291 metrics:292 - type: Accuracy293 value: 35.78294 - task:295 type: Natural language inference296 dataset:297 type: xnli298 name: XNLI (el)299 config: el300 split: validation301 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16302 metrics:303 - type: Accuracy304 value: 37.43305 - task:306 type: Natural language inference307 dataset:308 type: xnli309 name: XNLI (en)310 config: en311 split: validation312 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16313 metrics:314 - type: Accuracy315 value: 38.47316 - task:317 type: Natural language inference318 dataset:319 type: xnli320 name: XNLI (es)321 config: es322 split: validation323 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16324 metrics:325 - type: Accuracy326 value: 36.75327 - task:328 type: Natural language inference329 dataset:330 type: xnli331 name: XNLI (fr)332 config: fr333 split: validation334 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16335 metrics:336 - type: Accuracy337 value: 37.15338 - task:339 type: Natural language inference340 dataset:341 type: xnli342 name: XNLI (hi)343 config: hi344 split: validation345 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16346 metrics:347 - type: Accuracy348 value: 35.38349 - task:350 type: Natural language inference351 dataset:352 type: xnli353 name: XNLI (ru)354 config: ru355 split: validation356 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16357 metrics:358 - type: Accuracy359 value: 37.35360 - task:361 type: Natural language inference362 dataset:363 type: xnli364 name: XNLI (sw)365 config: sw366 split: validation367 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16368 metrics:369 - type: Accuracy370 value: 35.18371 - task:372 type: Natural language inference373 dataset:374 type: xnli375 name: XNLI (th)376 config: th377 split: validation378 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16379 metrics:380 - type: Accuracy381 value: 37.55382 - task:383 type: Natural language inference384 dataset:385 type: xnli386 name: XNLI (tr)387 config: tr388 split: validation389 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16390 metrics:391 - type: Accuracy392 value: 36.51393 - task:394 type: Natural language inference395 dataset:396 type: xnli397 name: XNLI (ur)398 config: ur399 split: validation400 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16401 metrics:402 - type: Accuracy403 value: 35.78404 - task:405 type: Natural language inference406 dataset:407 type: xnli408 name: XNLI (vi)409 config: vi410 split: validation411 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16412 metrics:413 - type: Accuracy414 value: 36.95415 - task:416 type: Natural language inference417 dataset:418 type: xnli419 name: XNLI (zh)420 config: zh421 split: validation422 revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16423 metrics:424 - type: Accuracy425 value: 37.07426 - task:427 type: Sentence completion428 dataset:429 type: story_cloze430 name: StoryCloze (2016)431 config: "2016"432 split: validation433 revision: e724c6f8cdf7c7a2fb229d862226e15b023ee4db434 metrics:435 - type: Accuracy436 value: 54.36437 - task:438 type: Sentence completion439 dataset:440 type: super_glue441 name: SuperGLUE (copa)442 config: copa443 split: validation444 revision: 9e12063561e7e6c79099feb6d5a493142584e9e2445 metrics:446 - type: Accuracy447 value: 57.0448 - task:449 type: Sentence completion450 dataset:451 type: xcopa452 name: XCOPA (et)453 config: et454 split: validation455 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187456 metrics:457 - type: Accuracy458 value: 57.0459 - task:460 type: Sentence completion461 dataset:462 type: xcopa463 name: XCOPA (ht)464 config: ht465 split: validation466 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187467 metrics:468 - type: Accuracy469 value: 60.0470 - task:471 type: Sentence completion472 dataset:473 type: xcopa474 name: XCOPA (id)475 config: id476 split: validation477 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187478 metrics:479 - type: Accuracy480 value: 59.0481 - task:482 type: Sentence completion483 dataset:484 type: xcopa485 name: XCOPA (it)486 config: it487 split: validation488 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187489 metrics:490 - type: Accuracy491 value: 59.0492 - task:493 type: Sentence completion494 dataset:495 type: xcopa496 name: XCOPA (qu)497 config: qu498 split: validation499 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187500 metrics:501 - type: Accuracy502 value: 54.0503 - task:504 type: Sentence completion505 dataset:506 type: xcopa507 name: XCOPA (sw)508 config: sw509 split: validation510 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187511 metrics:512 - type: Accuracy513 value: 55.0514 - task:515 type: Sentence completion516 dataset:517 type: xcopa518 name: XCOPA (ta)519 config: ta520 split: validation521 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187522 metrics:523 - type: Accuracy524 value: 59.0525 - task:526 type: Sentence completion527 dataset:528 type: xcopa529 name: XCOPA (th)530 config: th531 split: validation532 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187533 metrics:534 - type: Accuracy535 value: 65.0536 - task:537 type: Sentence completion538 dataset:539 type: xcopa540 name: XCOPA (tr)541 config: tr542 split: validation543 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187544 metrics:545 - type: Accuracy546 value: 58.0547 - task:548 type: Sentence completion549 dataset:550 type: xcopa551 name: XCOPA (vi)552 config: vi553 split: validation554 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187555 metrics:556 - type: Accuracy557 value: 54.0558 - task:559 type: Sentence completion560 dataset:561 type: xcopa562 name: XCOPA (zh)563 config: zh564 split: validation565 revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187566 metrics:567 - type: Accuracy568 value: 56.0569 - task:570 type: Sentence completion571 dataset:572 type: Muennighoff/xstory_cloze573 name: XStoryCloze (ar)574 config: ar575 split: validation576 revision: 8bb76e594b68147f1a430e86829d07189622b90d577 metrics:578 - type: Accuracy579 value: 48.78580 - task:581 type: Sentence completion582 dataset:583 type: Muennighoff/xstory_cloze584 name: XStoryCloze (es)585 config: es586 split: validation587 revision: 8bb76e594b68147f1a430e86829d07189622b90d588 metrics:589 - type: Accuracy590 value: 55.2591 - task:592 type: Sentence completion593 dataset:594 type: Muennighoff/xstory_cloze595 name: XStoryCloze (eu)596 config: eu597 split: validation598 revision: 8bb76e594b68147f1a430e86829d07189622b90d599 metrics:600 - type: Accuracy601 value: 52.95602 - task:603 type: Sentence completion604 dataset:605 type: Muennighoff/xstory_cloze606 name: XStoryCloze (hi)607 config: hi608 split: validation609 revision: 8bb76e594b68147f1a430e86829d07189622b90d610 metrics:611 - type: Accuracy612 value: 53.01613 - task:614 type: Sentence completion615 dataset:616 type: Muennighoff/xstory_cloze617 name: XStoryCloze (id)618 config: id619 split: validation620 revision: 8bb76e594b68147f1a430e86829d07189622b90d621 metrics:622 - type: Accuracy623 value: 53.08624 - task:625 type: Sentence completion626 dataset:627 type: Muennighoff/xstory_cloze628 name: XStoryCloze (my)629 config: my630 split: validation631 revision: 8bb76e594b68147f1a430e86829d07189622b90d632 metrics:633 - type: Accuracy634 value: 51.82635 - task:636 type: Sentence completion637 dataset:638 type: Muennighoff/xstory_cloze639 name: XStoryCloze (ru)640 config: ru641 split: validation642 revision: 8bb76e594b68147f1a430e86829d07189622b90d643 metrics:644 - type: Accuracy645 value: 49.7646 - task:647 type: Sentence completion648 dataset:649 type: Muennighoff/xstory_cloze650 name: XStoryCloze (sw)651 config: sw652 split: validation653 revision: 8bb76e594b68147f1a430e86829d07189622b90d654 metrics:655 - type: Accuracy656 value: 54.53657 - task:658 type: Sentence completion659 dataset:660 type: Muennighoff/xstory_cloze661 name: XStoryCloze (te)662 config: te663 split: validation664 revision: 8bb76e594b68147f1a430e86829d07189622b90d665 metrics:666 - type: Accuracy667 value: 53.67668 - task:669 type: Sentence completion670 dataset:671 type: Muennighoff/xstory_cloze672 name: XStoryCloze (zh)673 config: zh674 split: validation675 revision: 8bb76e594b68147f1a430e86829d07189622b90d676 metrics:677 - type: Accuracy678 value: 57.78679---680 681682 683# Table of Contents684 6851. [Model Summary](#model-summary)6862. [Use](#use)6873. [Limitations](#limitations)6884. [Training](#training)6895. [Evaluation](#evaluation)6907. [Citation](#citation)691 692# Model Summary693 694> We present BLOOMZ & mT0, a family of models capable of following human instructions in dozens of languages zero-shot. We finetune BLOOM & mT5 pretrained multilingual language models on our crosslingual task mixture (xP3) and find our resulting models capable of crosslingual generalization to unseen tasks & languages.695 696- **Repository:** [bigscience-workshop/xmtf](https://github.com/bigscience-workshop/xmtf)697- **Paper:** [Crosslingual Generalization through Multitask Finetuning](https://arxiv.org/abs/2211.01786)698- **Point of Contact:** [Niklas Muennighoff](mailto:niklas@hf.co)699- **Languages:** Refer to [mc4](https://huggingface.co/datasets/mc4) for pretraining & [xP3](https://huggingface.co/datasets/bigscience/xP3) for finetuning language proportions. It understands both pretraining & finetuning languages.700- **BLOOMZ & mT0 Model Family:**701 702<div class="max-w-full overflow-auto">703<table>704 <tr>705<th colspan="12">Multitask finetuned on <a style="font-weight:bold" href=https://huggingface.co/datasets/bigscience/xP3>xP3</a>. Recommended for prompting in English.706</tr>707<tr>708<td>Parameters</td>709<td>300M</td>710<td>580M</td>711<td>1.2B</td>712<td>3.7B</td>713<td>13B</td>714<td>560M</td>715<td>1.1B</td>716<td>1.7B</td>717<td>3B</td>718<td>7.1B</td>719<td>176B</td>720</tr>721<tr>722<td>Finetuned Model</td>723<td><a href=https://huggingface.co/bigscience/mt0-small>mt0-small</a></td> 724<td><a href=https://huggingface.co/bigscience/mt0-base>mt0-base</a></td>725<td><a href=https://huggingface.co/bigscience/mt0-large>mt0-large</a></td>726<td><a href=https://huggingface.co/bigscience/mt0-xl>mt0-xl</a></td>727<td><a href=https://huggingface.co/bigscience/mt0-xxl>mt0-xxl</a></td>728<td><a href=https://huggingface.co/bigscience/bloomz-560m>bloomz-560m</a></td>729<td><a href=https://huggingface.co/bigscience/bloomz-1b1>bloomz-1b1</a></td>730<td><a href=https://huggingface.co/bigscience/bloomz-1b7>bloomz-1b7</a></td>731<td><a href=https://huggingface.co/bigscience/bloomz-3b>bloomz-3b</a></td>732<td><a href=https://huggingface.co/bigscience/bloomz-7b1>bloomz-7b1</a></td>733<td><a href=https://huggingface.co/bigscience/bloomz>bloomz</a></td>734</tr>735</tr>736 <tr>737<th colspan="12">Multitask finetuned on <a style="font-weight:bold" href=https://huggingface.co/datasets/bigscience/xP3mt>xP3mt</a>. Recommended for prompting in non-English.</th>738</tr>739<tr>740<td>Finetuned Model</td>741<td></td>742<td></td>743<td></td>744<td></td>745<td><a href=https://huggingface.co/bigscience/mt0-xxl-mt>mt0-xxl-mt</a></td>746<td></td>747<td></td>748<td></td>749<td></td>750<td><a href=https://huggingface.co/bigscience/bloomz-7b1-mt>bloomz-7b1-mt</a></td>751<td><a href=https://huggingface.co/bigscience/bloomz-mt>bloomz-mt</a></td>752</tr>753<th colspan="12">Multitask finetuned on <a style="font-weight:bold" href=https://huggingface.co/datasets/Muennighoff/P3>P3</a>. Released for research purposes only. Strictly inferior to above models!</th>754</tr>755<tr>756<td>Finetuned Model</td>757<td></td>758<td></td>759<td></td>760<td></td>761<td><a href=https://huggingface.co/bigscience/mt0-xxl-p3>mt0-xxl-p3</a></td>762<td></td>763<td></td>764<td></td>765<td></td>766<td><a href=https://huggingface.co/bigscience/bloomz-7b1-p3>bloomz-7b1-p3</a></td>767<td><a href=https://huggingface.co/bigscience/bloomz-p3>bloomz-p3</a></td>768</tr>769<th colspan="12">Original pretrained checkpoints. Not recommended.</th>770<tr>771<td>Pretrained Model</td>772<td><a href=https://huggingface.co/google/mt5-small>mt5-small</a></td>773<td><a href=https://huggingface.co/google/mt5-base>mt5-base</a></td>774<td><a href=https://huggingface.co/google/mt5-large>mt5-large</a></td>775<td><a href=https://huggingface.co/google/mt5-xl>mt5-xl</a></td>776<td><a href=https://huggingface.co/google/mt5-xxl>mt5-xxl</a></td>777<td><a href=https://huggingface.co/bigscience/bloom-560m>bloom-560m</a></td>778<td><a href=https://huggingface.co/bigscience/bloom-1b1>bloom-1b1</a></td>779<td><a href=https://huggingface.co/bigscience/bloom-1b7>bloom-1b7</a></td>780<td><a href=https://huggingface.co/bigscience/bloom-3b>bloom-3b</a></td>781<td><a href=https://huggingface.co/bigscience/bloom-7b1>bloom-7b1</a></td>782<td><a href=https://huggingface.co/bigscience/bloom>bloom</a></td>783</tr>784</table>785</div>786 787 788# Use789 790## Intended use791 792We recommend using the model to perform tasks expressed in natural language. For example, given the prompt "*Translate to English: Je t’aime.*", the model will most likely answer "*I love you.*". Some prompt ideas from our paper: 793- 一个传奇的开端,一个不灭的神话,这不仅仅是一部电影,而是作为一个走进新时代的标签,永远彪炳史册。你认为这句话的立场是赞扬、中立还是批评?794- Suggest at least five related search terms to "Mạng neural nhân tạo".795- Write a fairy tale about a troll saving a princess from a dangerous dragon. The fairy tale is a masterpiece that has achieved praise worldwide and its moral is "Heroes Come in All Shapes and Sizes". Story (in Spanish):796- Explain in a sentence in Telugu what is backpropagation in neural networks.797 798**Feel free to share your generations in the Community tab!**799 800## How to use801 802### CPU803 804<details>805<summary> Click to expand </summary>806 807```python808# pip install -q transformers809from transformers import AutoModelForSeq2SeqLM, AutoTokenizer810 811checkpoint = "bigscience/mt0-small"812 813tokenizer = AutoTokenizer.from_pretrained(checkpoint)814model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint)815 816inputs = tokenizer.encode("Translate to English: Je t’aime.", return_tensors="pt")817outputs = model.generate(inputs)818print(tokenizer.decode(outputs[0]))819```820 821</details>822 823### GPU824 825<details>826<summary> Click to expand </summary>827 828```python829# pip install -q transformers accelerate830from transformers import AutoModelForSeq2SeqLM, AutoTokenizer831 832checkpoint = "bigscience/mt0-small"833 834tokenizer = AutoTokenizer.from_pretrained(checkpoint)835model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint, torch_dtype="auto", device_map="auto")836 837inputs = tokenizer.encode("Translate to English: Je t’aime.", return_tensors="pt").to("cuda")838outputs = model.generate(inputs)839print(tokenizer.decode(outputs[0]))840```841 842</details>843 844### GPU in 8bit845 846<details>847<summary> Click to expand </summary>848 849```python850# pip install -q transformers accelerate bitsandbytes851from transformers import AutoModelForSeq2SeqLM, AutoTokenizer852 853checkpoint = "bigscience/mt0-small"854 855tokenizer = AutoTokenizer.from_pretrained(checkpoint)856model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint, device_map="auto", load_in_8bit=True)857 858inputs = tokenizer.encode("Translate to English: Je t’aime.", return_tensors="pt").to("cuda")859outputs = model.generate(inputs)860print(tokenizer.decode(outputs[0]))861```862 863</details>864 865<!-- Necessary for whitespace -->866###867 868# Limitations869 870**Prompt Engineering:** The performance may vary depending on the prompt. For BLOOMZ models, we recommend making it very clear when the input stops to avoid the model trying to continue it. For example, the prompt "*Translate to English: Je t'aime*" without the full stop (.) at the end, may result in the model trying to continue the French sentence. Better prompts are e.g. "*Translate to English: Je t'aime.*", "*Translate to English: Je t'aime. Translation:*" "*What is "Je t'aime." in English?*", where it is clear for the model when it should answer. Further, we recommend providing the model as much context as possible. For example, if you want it to answer in Telugu, then tell the model, e.g. "*Explain in a sentence in Telugu what is backpropagation in neural networks.*".871 872# Training873 874## Model875 876- **Architecture:** Same as [mt5-small](https://huggingface.co/google/mt5-small), also refer to the `config.json` file877- **Finetuning steps:** 25000878- **Finetuning tokens:** 4.62 billion879- **Precision:** bfloat16880 881## Hardware882 883- **TPUs:** TPUv4-64884 885## Software886 887- **Orchestration:** [T5X](https://github.com/google-research/t5x)888- **Neural networks:** [Jax](https://github.com/google/jax)889 890# Evaluation891 892We refer to Table 7 from our [paper](https://arxiv.org/abs/2211.01786) & [bigscience/evaluation-results](https://huggingface.co/datasets/bigscience/evaluation-results) for zero-shot results on unseen tasks. The sidebar reports zero-shot performance of the best prompt per dataset config.893 894# Citation895```bibtex896@article{muennighoff2022crosslingual,897 title={Crosslingual generalization through multitask finetuning},898 author={Muennighoff, Niklas and Wang, Thomas and Sutawika, Lintang and Roberts, Adam and Biderman, Stella and Scao, Teven Le and Bari, M Saiful and Shen, Sheng and Yong, Zheng-Xin and Schoelkopf, Hailey and others},899 journal={arXiv preprint arXiv:2211.01786},900 year={2022}901}902```