CoolFace
Modelpublic

bigscience/mt0-small

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
32likes5.3kdownloads
README.md902 linesDownload Raw Back to root
1---2datasets:3- bigscience/xP34- mc45license: apache-2.06language:7- af8- am9- ar10- az11- be12- bg13- bn14- ca15- ceb16- co17- cs18- cy19- da20- de21- el22- en23- eo24- es25- et26- eu27- fa28- fi29- fil30- fr31- fy32- ga33- gd34- gl35- gu36- ha37- haw38- hi39- hmn40- ht41- hu42- hy43- ig44- is45- it46- iw47- ja48- jv49- ka50- kk51- km52- kn53- ko54- ku55- ky56- la57- lb58- lo59- lt60- lv61- mg62- mi63- mk64- ml65- mn66- mr67- ms68- mt69- my70- ne71- nl72- no73- ny74- pa75- pl76- ps77- pt78- ro79- ru80- sd81- si82- sk83- sl84- sm85- sn86- so87- sq88- sr89- st90- su91- sv92- sw93- ta94- te95- tg96- th97- tr98- uk99- und100- ur101- uz102- vi103- xh104- yi105- yo106- zh107- zu108pipeline_tag: text2text-generation109widget:110- text: "一个传奇的开端,一个不灭的神话,这不仅仅是一部电影,而是作为一个走进新时代的标签,永远彪炳史册。Would you rate the previous review as positive, neutral or negative?"111  example_title: "zh-en sentiment"112- text: "一个传奇的开端,一个不灭的神话,这不仅仅是一部电影,而是作为一个走进新时代的标签,永远彪炳史册。你认为这句话的立场是赞扬、中立还是批评?"113  example_title: "zh-zh sentiment"114- text: "Suggest at least five related search terms to \"Mạng neural nhân tạo\"."115  example_title: "vi-en query"116- text: "Proposez au moins cinq mots clés concernant «Réseau de neurones artificiels»."117  example_title: "fr-fr query"118- text: "Explain in a sentence in Telugu what is backpropagation in neural networks."119  example_title: "te-en qa"120- text: "Why is the sky blue?"121  example_title: "en-en qa"122- text: "Write a fairy tale about a troll saving a princess from a dangerous dragon. The fairy tale is a masterpiece that has achieved praise worldwide and its moral is \"Heroes Come in All Shapes and Sizes\". Story (in Spanish):"123  example_title: "es-en fable"124- text: "Write a fable about wood elves living in a forest that is suddenly invaded by ogres. The fable is a masterpiece that has achieved praise worldwide and its moral is \"Violence is the last refuge of the incompetent\". Fable (in Hindi):"125  example_title: "hi-en fable"126model-index:127- name: mt0-small128  results:129  - task:130      type: Coreference resolution131    dataset:132      type: winogrande133      name: Winogrande XL (xl)134      config: xl135      split: validation136      revision: a80f460359d1e9a67c006011c94de42a8759430c137    metrics:138    - type: Accuracy139      value: 50.51140  - task:141      type: Coreference resolution142    dataset:143      type: Muennighoff/xwinograd144      name: XWinograd (en)145      config: en146      split: test147      revision: 9dd5ea5505fad86b7bedad667955577815300cee148    metrics:149    - type: Accuracy150      value: 51.31151  - task:152      type: Coreference resolution153    dataset:154      type: Muennighoff/xwinograd155      name: XWinograd (fr)156      config: fr157      split: test158      revision: 9dd5ea5505fad86b7bedad667955577815300cee159    metrics:160    - type: Accuracy161      value: 54.22162  - task:163      type: Coreference resolution164    dataset:165      type: Muennighoff/xwinograd166      name: XWinograd (jp)167      config: jp168      split: test169      revision: 9dd5ea5505fad86b7bedad667955577815300cee170    metrics:171    - type: Accuracy172      value: 52.45173  - task:174      type: Coreference resolution175    dataset:176      type: Muennighoff/xwinograd177      name: XWinograd (pt)178      config: pt179      split: test180      revision: 9dd5ea5505fad86b7bedad667955577815300cee181    metrics:182    - type: Accuracy183      value: 51.71184  - task:185      type: Coreference resolution186    dataset:187      type: Muennighoff/xwinograd188      name: XWinograd (ru)189      config: ru190      split: test191      revision: 9dd5ea5505fad86b7bedad667955577815300cee192    metrics:193    - type: Accuracy194      value: 54.29195  - task:196      type: Coreference resolution197    dataset:198      type: Muennighoff/xwinograd199      name: XWinograd (zh)200      config: zh201      split: test202      revision: 9dd5ea5505fad86b7bedad667955577815300cee203    metrics:204    - type: Accuracy205      value: 54.17206  - task:207      type: Natural language inference208    dataset:209      type: anli210      name: ANLI (r1)211      config: r1212      split: validation213      revision: 9dbd830a06fea8b1c49d6e5ef2004a08d9f45094214    metrics:215    - type: Accuracy216      value: 34.7217  - task:218      type: Natural language inference219    dataset:220      type: anli221      name: ANLI (r2)222      config: r2223      split: validation224      revision: 9dbd830a06fea8b1c49d6e5ef2004a08d9f45094225    metrics:226    - type: Accuracy227      value: 34.0228  - task:229      type: Natural language inference230    dataset:231      type: anli232      name: ANLI (r3)233      config: r3234      split: validation235      revision: 9dbd830a06fea8b1c49d6e5ef2004a08d9f45094236    metrics:237    - type: Accuracy238      value: 33.83239  - task:240      type: Natural language inference241    dataset:242      type: super_glue243      name: SuperGLUE (cb)244      config: cb245      split: validation246      revision: 9e12063561e7e6c79099feb6d5a493142584e9e2247    metrics:248    - type: Accuracy249      value: 50.0250  - task:251      type: Natural language inference252    dataset:253      type: super_glue254      name: SuperGLUE (rte)255      config: rte256      split: validation257      revision: 9e12063561e7e6c79099feb6d5a493142584e9e2258    metrics:259    - type: Accuracy260      value: 61.01261  - task:262      type: Natural language inference263    dataset:264      type: xnli265      name: XNLI (ar)266      config: ar267      split: validation268      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16269    metrics:270    - type: Accuracy271      value: 37.43272  - task:273      type: Natural language inference274    dataset:275      type: xnli276      name: XNLI (bg)277      config: bg278      split: validation279      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16280    metrics:281    - type: Accuracy282      value: 37.55283  - task:284      type: Natural language inference285    dataset:286      type: xnli287      name: XNLI (de)288      config: de289      split: validation290      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16291    metrics:292    - type: Accuracy293      value: 35.78294  - task:295      type: Natural language inference296    dataset:297      type: xnli298      name: XNLI (el)299      config: el300      split: validation301      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16302    metrics:303    - type: Accuracy304      value: 37.43305  - task:306      type: Natural language inference307    dataset:308      type: xnli309      name: XNLI (en)310      config: en311      split: validation312      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16313    metrics:314    - type: Accuracy315      value: 38.47316  - task:317      type: Natural language inference318    dataset:319      type: xnli320      name: XNLI (es)321      config: es322      split: validation323      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16324    metrics:325    - type: Accuracy326      value: 36.75327  - task:328      type: Natural language inference329    dataset:330      type: xnli331      name: XNLI (fr)332      config: fr333      split: validation334      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16335    metrics:336    - type: Accuracy337      value: 37.15338  - task:339      type: Natural language inference340    dataset:341      type: xnli342      name: XNLI (hi)343      config: hi344      split: validation345      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16346    metrics:347    - type: Accuracy348      value: 35.38349  - task:350      type: Natural language inference351    dataset:352      type: xnli353      name: XNLI (ru)354      config: ru355      split: validation356      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16357    metrics:358    - type: Accuracy359      value: 37.35360  - task:361      type: Natural language inference362    dataset:363      type: xnli364      name: XNLI (sw)365      config: sw366      split: validation367      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16368    metrics:369    - type: Accuracy370      value: 35.18371  - task:372      type: Natural language inference373    dataset:374      type: xnli375      name: XNLI (th)376      config: th377      split: validation378      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16379    metrics:380    - type: Accuracy381      value: 37.55382  - task:383      type: Natural language inference384    dataset:385      type: xnli386      name: XNLI (tr)387      config: tr388      split: validation389      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16390    metrics:391    - type: Accuracy392      value: 36.51393  - task:394      type: Natural language inference395    dataset:396      type: xnli397      name: XNLI (ur)398      config: ur399      split: validation400      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16401    metrics:402    - type: Accuracy403      value: 35.78404  - task:405      type: Natural language inference406    dataset:407      type: xnli408      name: XNLI (vi)409      config: vi410      split: validation411      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16412    metrics:413    - type: Accuracy414      value: 36.95415  - task:416      type: Natural language inference417    dataset:418      type: xnli419      name: XNLI (zh)420      config: zh421      split: validation422      revision: a5a45e4ff92d5d3f34de70aaf4b72c3bdf9f7f16423    metrics:424    - type: Accuracy425      value: 37.07426  - task:427      type: Sentence completion428    dataset:429      type: story_cloze430      name: StoryCloze (2016)431      config: "2016"432      split: validation433      revision: e724c6f8cdf7c7a2fb229d862226e15b023ee4db434    metrics:435    - type: Accuracy436      value: 54.36437  - task:438      type: Sentence completion439    dataset:440      type: super_glue441      name: SuperGLUE (copa)442      config: copa443      split: validation444      revision: 9e12063561e7e6c79099feb6d5a493142584e9e2445    metrics:446    - type: Accuracy447      value: 57.0448  - task:449      type: Sentence completion450    dataset:451      type: xcopa452      name: XCOPA (et)453      config: et454      split: validation455      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187456    metrics:457    - type: Accuracy458      value: 57.0459  - task:460      type: Sentence completion461    dataset:462      type: xcopa463      name: XCOPA (ht)464      config: ht465      split: validation466      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187467    metrics:468    - type: Accuracy469      value: 60.0470  - task:471      type: Sentence completion472    dataset:473      type: xcopa474      name: XCOPA (id)475      config: id476      split: validation477      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187478    metrics:479    - type: Accuracy480      value: 59.0481  - task:482      type: Sentence completion483    dataset:484      type: xcopa485      name: XCOPA (it)486      config: it487      split: validation488      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187489    metrics:490    - type: Accuracy491      value: 59.0492  - task:493      type: Sentence completion494    dataset:495      type: xcopa496      name: XCOPA (qu)497      config: qu498      split: validation499      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187500    metrics:501    - type: Accuracy502      value: 54.0503  - task:504      type: Sentence completion505    dataset:506      type: xcopa507      name: XCOPA (sw)508      config: sw509      split: validation510      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187511    metrics:512    - type: Accuracy513      value: 55.0514  - task:515      type: Sentence completion516    dataset:517      type: xcopa518      name: XCOPA (ta)519      config: ta520      split: validation521      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187522    metrics:523    - type: Accuracy524      value: 59.0525  - task:526      type: Sentence completion527    dataset:528      type: xcopa529      name: XCOPA (th)530      config: th531      split: validation532      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187533    metrics:534    - type: Accuracy535      value: 65.0536  - task:537      type: Sentence completion538    dataset:539      type: xcopa540      name: XCOPA (tr)541      config: tr542      split: validation543      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187544    metrics:545    - type: Accuracy546      value: 58.0547  - task:548      type: Sentence completion549    dataset:550      type: xcopa551      name: XCOPA (vi)552      config: vi553      split: validation554      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187555    metrics:556    - type: Accuracy557      value: 54.0558  - task:559      type: Sentence completion560    dataset:561      type: xcopa562      name: XCOPA (zh)563      config: zh564      split: validation565      revision: 37f73c60fb123111fa5af5f9b705d0b3747fd187566    metrics:567    - type: Accuracy568      value: 56.0569  - task:570      type: Sentence completion571    dataset:572      type: Muennighoff/xstory_cloze573      name: XStoryCloze (ar)574      config: ar575      split: validation576      revision: 8bb76e594b68147f1a430e86829d07189622b90d577    metrics:578    - type: Accuracy579      value: 48.78580  - task:581      type: Sentence completion582    dataset:583      type: Muennighoff/xstory_cloze584      name: XStoryCloze (es)585      config: es586      split: validation587      revision: 8bb76e594b68147f1a430e86829d07189622b90d588    metrics:589    - type: Accuracy590      value: 55.2591  - task:592      type: Sentence completion593    dataset:594      type: Muennighoff/xstory_cloze595      name: XStoryCloze (eu)596      config: eu597      split: validation598      revision: 8bb76e594b68147f1a430e86829d07189622b90d599    metrics:600    - type: Accuracy601      value: 52.95602  - task:603      type: Sentence completion604    dataset:605      type: Muennighoff/xstory_cloze606      name: XStoryCloze (hi)607      config: hi608      split: validation609      revision: 8bb76e594b68147f1a430e86829d07189622b90d610    metrics:611    - type: Accuracy612      value: 53.01613  - task:614      type: Sentence completion615    dataset:616      type: Muennighoff/xstory_cloze617      name: XStoryCloze (id)618      config: id619      split: validation620      revision: 8bb76e594b68147f1a430e86829d07189622b90d621    metrics:622    - type: Accuracy623      value: 53.08624  - task:625      type: Sentence completion626    dataset:627      type: Muennighoff/xstory_cloze628      name: XStoryCloze (my)629      config: my630      split: validation631      revision: 8bb76e594b68147f1a430e86829d07189622b90d632    metrics:633    - type: Accuracy634      value: 51.82635  - task:636      type: Sentence completion637    dataset:638      type: Muennighoff/xstory_cloze639      name: XStoryCloze (ru)640      config: ru641      split: validation642      revision: 8bb76e594b68147f1a430e86829d07189622b90d643    metrics:644    - type: Accuracy645      value: 49.7646  - task:647      type: Sentence completion648    dataset:649      type: Muennighoff/xstory_cloze650      name: XStoryCloze (sw)651      config: sw652      split: validation653      revision: 8bb76e594b68147f1a430e86829d07189622b90d654    metrics:655    - type: Accuracy656      value: 54.53657  - task:658      type: Sentence completion659    dataset:660      type: Muennighoff/xstory_cloze661      name: XStoryCloze (te)662      config: te663      split: validation664      revision: 8bb76e594b68147f1a430e86829d07189622b90d665    metrics:666    - type: Accuracy667      value: 53.67668  - task:669      type: Sentence completion670    dataset:671      type: Muennighoff/xstory_cloze672      name: XStoryCloze (zh)673      config: zh674      split: validation675      revision: 8bb76e594b68147f1a430e86829d07189622b90d676    metrics:677    - type: Accuracy678      value: 57.78679---680 681![xmtf](https://github.com/bigscience-workshop/xmtf/blob/master/xmtf_banner.png?raw=true)682 683#  Table of Contents684 6851. [Model Summary](#model-summary)6862. [Use](#use)6873. [Limitations](#limitations)6884. [Training](#training)6895. [Evaluation](#evaluation)6907. [Citation](#citation)691 692# Model Summary693 694> We present BLOOMZ & mT0, a family of models capable of following human instructions in dozens of languages zero-shot. We finetune BLOOM & mT5 pretrained multilingual language models on our crosslingual task mixture (xP3) and find our resulting models capable of crosslingual generalization to unseen tasks & languages.695 696- **Repository:** [bigscience-workshop/xmtf](https://github.com/bigscience-workshop/xmtf)697- **Paper:** [Crosslingual Generalization through Multitask Finetuning](https://arxiv.org/abs/2211.01786)698- **Point of Contact:** [Niklas Muennighoff](mailto:niklas@hf.co)699- **Languages:** Refer to [mc4](https://huggingface.co/datasets/mc4) for pretraining & [xP3](https://huggingface.co/datasets/bigscience/xP3) for finetuning language proportions. It understands both pretraining & finetuning languages.700- **BLOOMZ & mT0 Model Family:**701 702<div class="max-w-full overflow-auto">703<table>704  <tr>705<th colspan="12">Multitask finetuned on <a style="font-weight:bold" href=https://huggingface.co/datasets/bigscience/xP3>xP3</a>. Recommended for prompting in English.706</tr>707<tr>708<td>Parameters</td>709<td>300M</td>710<td>580M</td>711<td>1.2B</td>712<td>3.7B</td>713<td>13B</td>714<td>560M</td>715<td>1.1B</td>716<td>1.7B</td>717<td>3B</td>718<td>7.1B</td>719<td>176B</td>720</tr>721<tr>722<td>Finetuned Model</td>723<td><a href=https://huggingface.co/bigscience/mt0-small>mt0-small</a></td>  724<td><a href=https://huggingface.co/bigscience/mt0-base>mt0-base</a></td>725<td><a href=https://huggingface.co/bigscience/mt0-large>mt0-large</a></td>726<td><a href=https://huggingface.co/bigscience/mt0-xl>mt0-xl</a></td>727<td><a href=https://huggingface.co/bigscience/mt0-xxl>mt0-xxl</a></td>728<td><a href=https://huggingface.co/bigscience/bloomz-560m>bloomz-560m</a></td>729<td><a href=https://huggingface.co/bigscience/bloomz-1b1>bloomz-1b1</a></td>730<td><a href=https://huggingface.co/bigscience/bloomz-1b7>bloomz-1b7</a></td>731<td><a href=https://huggingface.co/bigscience/bloomz-3b>bloomz-3b</a></td>732<td><a href=https://huggingface.co/bigscience/bloomz-7b1>bloomz-7b1</a></td>733<td><a href=https://huggingface.co/bigscience/bloomz>bloomz</a></td>734</tr>735</tr>736  <tr>737<th colspan="12">Multitask finetuned on <a style="font-weight:bold" href=https://huggingface.co/datasets/bigscience/xP3mt>xP3mt</a>. Recommended for prompting in non-English.</th>738</tr>739<tr>740<td>Finetuned Model</td>741<td></td>742<td></td>743<td></td>744<td></td>745<td><a href=https://huggingface.co/bigscience/mt0-xxl-mt>mt0-xxl-mt</a></td>746<td></td>747<td></td>748<td></td>749<td></td>750<td><a href=https://huggingface.co/bigscience/bloomz-7b1-mt>bloomz-7b1-mt</a></td>751<td><a href=https://huggingface.co/bigscience/bloomz-mt>bloomz-mt</a></td>752</tr>753<th colspan="12">Multitask finetuned on <a style="font-weight:bold" href=https://huggingface.co/datasets/Muennighoff/P3>P3</a>. Released for research purposes only. Strictly inferior to above models!</th>754</tr>755<tr>756<td>Finetuned Model</td>757<td></td>758<td></td>759<td></td>760<td></td>761<td><a href=https://huggingface.co/bigscience/mt0-xxl-p3>mt0-xxl-p3</a></td>762<td></td>763<td></td>764<td></td>765<td></td>766<td><a href=https://huggingface.co/bigscience/bloomz-7b1-p3>bloomz-7b1-p3</a></td>767<td><a href=https://huggingface.co/bigscience/bloomz-p3>bloomz-p3</a></td>768</tr>769<th colspan="12">Original pretrained checkpoints. Not recommended.</th>770<tr>771<td>Pretrained Model</td>772<td><a href=https://huggingface.co/google/mt5-small>mt5-small</a></td>773<td><a href=https://huggingface.co/google/mt5-base>mt5-base</a></td>774<td><a href=https://huggingface.co/google/mt5-large>mt5-large</a></td>775<td><a href=https://huggingface.co/google/mt5-xl>mt5-xl</a></td>776<td><a href=https://huggingface.co/google/mt5-xxl>mt5-xxl</a></td>777<td><a href=https://huggingface.co/bigscience/bloom-560m>bloom-560m</a></td>778<td><a href=https://huggingface.co/bigscience/bloom-1b1>bloom-1b1</a></td>779<td><a href=https://huggingface.co/bigscience/bloom-1b7>bloom-1b7</a></td>780<td><a href=https://huggingface.co/bigscience/bloom-3b>bloom-3b</a></td>781<td><a href=https://huggingface.co/bigscience/bloom-7b1>bloom-7b1</a></td>782<td><a href=https://huggingface.co/bigscience/bloom>bloom</a></td>783</tr>784</table>785</div>786 787 788# Use789 790## Intended use791 792We recommend using the model to perform tasks expressed in natural language. For example, given the prompt "*Translate to English: Je t’aime.*", the model will most likely answer "*I love you.*". Some prompt ideas from our paper: 793- 一个传奇的开端,一个不灭的神话,这不仅仅是一部电影,而是作为一个走进新时代的标签,永远彪炳史册。你认为这句话的立场是赞扬、中立还是批评?794- Suggest at least five related search terms to "Mạng neural nhân tạo".795- Write a fairy tale about a troll saving a princess from a dangerous dragon. The fairy tale is a masterpiece that has achieved praise worldwide and its moral is "Heroes Come in All Shapes and Sizes". Story (in Spanish):796- Explain in a sentence in Telugu what is backpropagation in neural networks.797 798**Feel free to share your generations in the Community tab!**799 800## How to use801 802### CPU803 804<details>805<summary> Click to expand </summary>806 807```python808# pip install -q transformers809from transformers import AutoModelForSeq2SeqLM, AutoTokenizer810 811checkpoint = "bigscience/mt0-small"812 813tokenizer = AutoTokenizer.from_pretrained(checkpoint)814model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint)815 816inputs = tokenizer.encode("Translate to English: Je t’aime.", return_tensors="pt")817outputs = model.generate(inputs)818print(tokenizer.decode(outputs[0]))819```820 821</details>822 823### GPU824 825<details>826<summary> Click to expand </summary>827 828```python829# pip install -q transformers accelerate830from transformers import AutoModelForSeq2SeqLM, AutoTokenizer831 832checkpoint = "bigscience/mt0-small"833 834tokenizer = AutoTokenizer.from_pretrained(checkpoint)835model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint, torch_dtype="auto", device_map="auto")836 837inputs = tokenizer.encode("Translate to English: Je t’aime.", return_tensors="pt").to("cuda")838outputs = model.generate(inputs)839print(tokenizer.decode(outputs[0]))840```841 842</details>843 844### GPU in 8bit845 846<details>847<summary> Click to expand </summary>848 849```python850# pip install -q transformers accelerate bitsandbytes851from transformers import AutoModelForSeq2SeqLM, AutoTokenizer852 853checkpoint = "bigscience/mt0-small"854 855tokenizer = AutoTokenizer.from_pretrained(checkpoint)856model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint, device_map="auto", load_in_8bit=True)857 858inputs = tokenizer.encode("Translate to English: Je t’aime.", return_tensors="pt").to("cuda")859outputs = model.generate(inputs)860print(tokenizer.decode(outputs[0]))861```862 863</details>864 865<!-- Necessary for whitespace -->866###867 868# Limitations869 870**Prompt Engineering:** The performance may vary depending on the prompt. For BLOOMZ models, we recommend making it very clear when the input stops to avoid the model trying to continue it. For example, the prompt "*Translate to English: Je t'aime*" without the full stop (.) at the end, may result in the model trying to continue the French sentence. Better prompts are e.g. "*Translate to English: Je t'aime.*", "*Translate to English: Je t'aime. Translation:*" "*What is "Je t'aime." in English?*", where it is clear for the model when it should answer. Further, we recommend providing the model as much context as possible. For example, if you want it to answer in Telugu, then tell the model, e.g. "*Explain in a sentence in Telugu what is backpropagation in neural networks.*".871 872# Training873 874## Model875 876- **Architecture:** Same as [mt5-small](https://huggingface.co/google/mt5-small), also refer to the `config.json` file877- **Finetuning steps:** 25000878- **Finetuning tokens:** 4.62 billion879- **Precision:** bfloat16880 881## Hardware882 883- **TPUs:** TPUv4-64884 885## Software886 887- **Orchestration:** [T5X](https://github.com/google-research/t5x)888- **Neural networks:** [Jax](https://github.com/google/jax)889 890# Evaluation891 892We refer to Table 7 from our [paper](https://arxiv.org/abs/2211.01786) & [bigscience/evaluation-results](https://huggingface.co/datasets/bigscience/evaluation-results) for zero-shot results on unseen tasks. The sidebar reports zero-shot performance of the best prompt per dataset config.893 894# Citation895```bibtex896@article{muennighoff2022crosslingual,897  title={Crosslingual generalization through multitask finetuning},898  author={Muennighoff, Niklas and Wang, Thomas and Sutawika, Lintang and Roberts, Adam and Biderman, Stella and Scao, Teven Le and Bari, M Saiful and Shen, Sheng and Yong, Zheng-Xin and Schoelkopf, Hailey and others},899  journal={arXiv preprint arXiv:2211.01786},900  year={2022}901}902```