CoolFace
Modelpublic

multimolecule/splicebert.510

sourceHugging Faceagpl-3.0updated 4mo agoView on Hugging Face
0likes67downloads
README.md819 linesDownload Raw Back to root
1---2datasets:3- multimolecule/ucsc-genome-browser4library_name: multimolecule5license: agpl-3.06mask_token: <mask>7pipeline_tag: fill-mask8tags:9- Biology10- RNA11- ncRNA12- rna13widget:14- example_title: microRNA 2115  mask_index: 1116  mask_index_1based: 1217  masked_char: A18  output:19  - label: C20    score: 0.07627921  - label: S22    score: 0.06061423  - label: M24    score: 0.05649125  - label: V26    score: 0.05356727  - label: Y28    score: 0.04820329  pipeline_tag: fill-mask30  sequence_type: ncRNA31  task: fill-mask32  text: UAGCUUAUCAG<mask>CUGAUGUUGA33- example_title: microRNA 146a34  mask_index: 1035  mask_index_1based: 1136  masked_char: A37  output:38  - label: G39    score: 0.16400540  - label: R41    score: 0.08606942  - label: K43    score: 0.06784844  - label: D45    score: 0.05924346  - label: A47    score: 0.04516948  pipeline_tag: fill-mask49  sequence_type: ncRNA50  task: fill-mask51  text: UGAGAACUGA<mask>UUCCAUGGGUU52- example_title: microRNA 15553  mask_index: 1554  mask_index_1based: 1655  masked_char: A56  output:57  - label: U58    score: 0.06264659  - label: W60    score: 0.05326461  - label: K62    score: 0.05004963  - label: Y64    score: 0.04972265  - label: D66    score: 0.04840867  pipeline_tag: fill-mask68  sequence_type: ncRNA69  task: fill-mask70  text: UUAAUGCUAAUCGUG<mask>UAGGGGUU71- example_title: RNA component of mitochondrial RNA processing endoribonuclease72  mask_index: 1173  mask_index_1based: 1274  masked_char: A75  output:76  - label: C77    score: 0.07944578  - label: S79    score: 0.06397580  - label: M81    score: 0.05370982  - label: V83    score: 0.05296984  - label: G85    score: 0.05151886  pipeline_tag: fill-mask87  sequence_type: ncRNA88  task: fill-mask89  text: GGUUCGUGCUG<mask>AGGCCUGUAUCCUAGGCUACACACUGAGGACUCUGUUCCUCCCCUUUCCGCCUAGGGGAAAGUCCCCGGACCUCGGGCAGAGAGUGCCACGUGCAUACGCACGUAGACAUUCCCCGCUUCCCACUCCAAAGUCCGCCAAGAAGCGUAUCCCGCUGAGCGGCGUGGCGCGGGGGCGUCAUCCGUCAGCUCCCUCUAGUUACGCAGGCAGUGCGUGUCCGCGCACCAACCACACGGGGCUCAUUCUCAGCGCGGCUGUAAAAAAAAA90- example_title: 7SK small nuclear RNA91  mask_index: 1392  mask_index_1based: 1493  masked_char: A94  output:95  - label: A96    score: 0.06808697  - label: R98    score: 0.06639399  - label: G100    score: 0.064741101  - label: V102    score: 0.057095103  - label: M104    score: 0.053618105  pipeline_tag: fill-mask106  sequence_type: ncRNA107  task: fill-mask108  text: GGAUGUGAGGGCG<mask>UCUGGCUGCGACAUCUGUCACCCCAUUGAUCGCCAGGGUUGAUUCGGCUGAUCUGGCUGGCUAGGCGGGUGUCCCCUUCCUCCCUCACCGCUCCAUGUGCGUCCCUCCCGAAGCUGCGCGCUCGGUCGAAGAGGACGACCAUCCCCGAUAGAGGAGGACCGGUCUUCGGUCAAGGGUAUACGAGUAGCUGCGCUCCCCUGCUAGAACCUCCAAACAAGCUCUCAAGGUCCAUUUGUAGGAGAACGUAGGGUAGUCAAGCUUCCAAGACUCCAGACACAUCCAAAUGAGGCGCUGCAUGUGGCAGUCUGCCUUUCUUUU109- example_title: telomerase RNA component110  mask_index: 23111  mask_index_1based: 24112  masked_char: A113  output:114  - label: C115    score: 0.061932116  - label: Y117    score: 0.061489118  - label: U119    score: 0.061049120  - label: H121    score: 0.060014122  - label: M123    score: 0.059503124  pipeline_tag: fill-mask125  sequence_type: ncRNA126  task: fill-mask127  text: GGGUUGCGGAGGGUGGGCCUGGG<mask>GGGGUGGUGGCCAUUUUUUGUCUAACCCUAACUGAGAAGGGCGUAGGCGCCGUGCUUUUGCUCCCCGCGCGCUGUUUUUCUCGCUGACUUUCAGCGGGCGGAAAAGCCUCGGCCUGCCGCCUUCCACCGUUCAUUCUAGAGCAAACAAAAAAUGUCAGCUGCUGGCCCGUUCGCCCCUCCCGGGGACCUGCGGCGGGUCGCCUGCCCAGCCCCCGAACCCCGCCUGGAGGCCGCGGUCGGCCCGGGGCUUCUCCGGAGGCACCCACUGCCACCGCGAAGAGUUGGGCUCUGUCAGCCGCGGGUCUCUCGGGGGCGAGGGCGAGGUUCAGGCCUUUCAGGCCGCAGGAAGAGGAACGGAGCGAGUCCCCGCGCGCGGCGCGAUUCCCUGAGCUGUGGGACGUGCACCCAGGACUCGGCUCACACAUGC128- example_title: vault RNA 2-1129  mask_index: 12130  mask_index_1based: 13131  masked_char: A132  output:133  - label: U134    score: 0.074338135  - label: K136    score: 0.062814137  - label: Y138    score: 0.053943139  - label: B140    score: 0.053653141  - label: G142    score: 0.053077143  pipeline_tag: fill-mask144  sequence_type: ncRNA145  task: fill-mask146  text: CGGGUCGGAGUU<mask>GCUCAAGCGGUUACCUCCUCAUGCCGGACUUUCUAUCUGUCCAUCUCUGUGCUGGGGUUCGAGACCCGCGGGUGCUUACUGACCCUUUUAUGCAA147- example_title: brain cytoplasmic RNA 1148  mask_index: 18149  mask_index_1based: 19150  masked_char: A151  output:152  - label: A153    score: 0.416301154  - label: R155    score: 0.08191156  - label: M157    score: 0.060285158  - label: W159    score: 0.052732160  - label: I161    score: 0.039218162  pipeline_tag: fill-mask163  sequence_type: ncRNA164  task: fill-mask165  text: GGCCGGGCGCGGUGGCUC<mask>CGCCUGUAAUCCCAGCUCUCAGGGAGGCUAAGAGGCGGGAGGAUAGCUUGAGCCCAGGAGUUCGAGACCUGCCUGGGCAAUAUAGCGAGACCCCGUUCUCCAGAAAAAGGAAAAAAAAAAACAAAAGACAAAAAAAAAAUAAGCGUAACUUCCCUCAAAGCAACAACCCCCCCCCCCCUUU166- example_title: HIV-1 TAR-WT167  mask_index: 13168  mask_index_1based: 14169  masked_char: A170  output:171  - label: A172    score: 0.087018173  - label: R174    score: 0.078768175  - label: W176    score: 0.075575177  - label: D178    score: 0.074122179  - label: G180    score: 0.0713181  pipeline_tag: fill-mask182  sequence_type: ncRNA183  task: fill-mask184  text: GGUCUCUCUGGUU<mask>GACCAGAUCUGAGCCUGGGAGCUCUCUGGCUAACUAGGGAACC185- example_title: prion protein (Kanno blood group)186  mask_index: 21187  mask_index_1based: 22188  masked_char: A189  output:190  - label: C191    score: 0.203007192  - label: S193    score: 0.088737194  - label: Y195    score: 0.059722196  - label: M197    score: 0.0563198  - label: B199    score: 0.051719200  pipeline_tag: fill-mask201  sequence_type: mRNA202  task: fill-mask203  text: AUGGCGAACCUUGGCUGCUGG<mask>UGCUGGUUCUCUUUGUGGCCACAUGGAGUGACCUGGGCCUCUGC204- example_title: interleukin 10205  mask_index: 11206  mask_index_1based: 12207  masked_char: A208  output:209  - label: U210    score: 0.145719211  - label: W212    score: 0.092637213  - label: A214    score: 0.058892215  - label: H216    score: 0.054023217  - label: Y218    score: 0.051742219  pipeline_tag: fill-mask220  sequence_type: mRNA221  task: fill-mask222  text: AUGCACAGCUC<mask>GCACUGCUCUGUUGCCUGGUCCUCCUGACUGGGGUGAGGGCC223- example_title: Zaire ebolavirus224  mask_index: 11225  mask_index_1based: 12226  masked_char: A227  output:228  - label: U229    score: 0.089489230  - label: W231    score: 0.088384232  - label: A233    score: 0.087294234  - label: H235    score: 0.065665236  - label: Y237    score: 0.056951238  pipeline_tag: fill-mask239  sequence_type: mRNA240  task: fill-mask241  text: AAUGUUCAAAC<mask>CUUUGUGAAGCUCUGUUAGCUGAUGGUCUUGCUAAAGCAUUUCCUAGCAAUAUGAUGGUAGUCACAGAGCGUGAGCAAAAAGAAAGCUUAUUGCAUCAAGCAUCAUGGCACCACACAAGUGAUGAUUUUGGUGAGCAUGCCACAGUUAGAGGGAGUAGCUUUGUAACUGAUUUAGAGAAAUACAAUCUUGCAUUUAGAUAUGAGUUUACAGCACCUUUUAUAGAAUAUUGUAACCGUUGCUAUGGUGUUAAGAAUGUUUUUAAUUGGAUGCAUUAUACAAUCCCACAGUGUUAU242- example_title: SARS coronavirus243  mask_index: 14244  mask_index_1based: 15245  masked_char: A246  output:247  - label: U248    score: 0.124725249  - label: Y250    score: 0.065713251  - label: W252    score: 0.065298253  - label: K254    score: 0.057034255  - label: H256    score: 0.05285257  pipeline_tag: fill-mask258  sequence_type: mRNA259  task: fill-mask260  text: AUGUUUAUUUUCUU<mask>UUAUUUCUUACUCUCACUAGUGGUAGUGACCUUGACCGGUGCACCACUUUUGAUGAUGUUCAAGCUCCUAAUUACACUCAACAUACUUCAUCUAUGAGGGGGGUUUACUAUCCUGAUGAAAUUUUUAGAUCAGACACUCUUUAUUUAACUCAGGAUUUAUUUCUUCCAUUUUAUUCUAAUGUUACAGGGUUUCAUACUAUUAAUCAUACGUUUGACAACCCUGUCAUACCUUUUAAGGAUGGUAUUUAUUUUGCUGCCACAGAGAAAUCAAAUGUUGUCCGUGGUUGGGUUUUUGGUUCUACCAUGAACAACAAGUCACAGUCGGUGAUUAUUAUUAACAAUUCUACUAAUGUUGUUAUACGAGCAUGUAACUUUGAAUUGUGUGACAACCCUUUCUUUGCUGUUUCUAAACCCAUGGGUACACAGACACAUACUAUGAUAUUCGAUAAUGCAUUUAAAUGCACUUUCGAGUACAUAUCU261- example_title: insulin262  mask_index: 12263  mask_index_1based: 13264  masked_char: A265  output:266  - label: C267    score: 0.280004268  - label: S269    score: 0.107256270  - label: M271    score: 0.057172272  - label: Y273    score: 0.053979274  - label: V275    score: 0.05121276  pipeline_tag: fill-mask277  sequence_type: mRNA278  task: fill-mask279  text: AUGGCCCUGUGG<mask>UGCGCCUCCUGCCCCUGCUGGCGCUGCUGGCCCUCUGGGGACCUGACCCAGCCGCAGCCUUUGUGAACCAACACCUGUGCGGCUCACACCUGGUGGAAGCUCUCUACCUAGUGUGCGGGGAACGAGGCUUCUUCUACACACCCAAGACCCGCCGGGAGGCAGAGGACCUGCAGGUGGGGCAGGUGGAGCUGGGCGGGGGCCCUGGUGCAGGCAGCCUGCAGCCCUUGGCCCUGGAGGGGUCCCUGCAGAAGCGUGGCAUUGUGGAACAAUGCUGUACCAGCAUCUGCUCCCUCUACCAGCUGGAGAACUACUGCAACUAG280- example_title: cyclin dependent kinase inhibitor 2A281  mask_index: 18282  mask_index_1based: 19283  masked_char: A284  output:285  - label: G286    score: 0.098564287  - label: S288    score: 0.08414289  - label: C290    score: 0.071827291  - label: V292    score: 0.066422293  - label: R294    score: 0.063874295  pipeline_tag: fill-mask296  sequence_type: mRNA297  task: fill-mask298  text: AUGGAGCCGGCGGCGGGG<mask>GCAGCAUGGAGCCUUCGGCUGACUGGCUGGCCACGGCCGCGGCCCGGGGUCGGGUAGAGGAGGUGCGGGCGCUGCUGGAGGCGGGGGCGCUGCCCAACGCACCGAAUAGUUACGGUCGGAGGCCGAUCCAGGUCAUGAUGAUGGGCAGCGCCCGAGUGGCGGAGCUGCUGCUGCUCCACGGCGCGGAGCCCAACUGCGCCGACCCCGCCACUCUCACCCGACCCGUGCACGACGCUGCCCGGGAGGGCUUCCUGGACACGCUGGUGGUGCUGCACCGGGCCGGGGCGCGGCUGGACGUGCGCGAUGCCUGGGGCCGUCUGCCCGUGGACCUGGCUGAGGAGCUGGGCCAUCGCGAUGUCGCACGGUACCUGCGCGCGGCUGCGGGGGGCACCAGAGGCAGUAACCAUGCCCGCAUAGAUGCCGCGGAAGGUCCCUCAGACAUCCCCGAUUGA299- example_title: human papillomavirus type 16 E6300  mask_index: 10301  mask_index_1based: 11302  masked_char: A303  output:304  - label: A305    score: 0.086406306  - label: W307    score: 0.077731308  - label: U309    score: 0.069927310  - label: D311    score: 0.068029312  - label: R313    score: 0.0671314  pipeline_tag: fill-mask315  sequence_type: mRNA316  task: fill-mask317  text: AUGCACCAAA<mask>GAGAACUGCAAUGUUUCAGGACCCACAGGAGCGACCCAGAAAGUUACCACAGUUAUGCACAGAGCUGCAAACAACUAUACAUGAUAUAAUAUUAGAAUGUGUGUACUGCAAGCAACAGUUACUGCGACGUGAGGUAUAUGACUUUGCUUUUCGGGAUUUAUGCAUAGUAUAUAGAGAUGGGAAUCCAUAUGCUGUAUGUGAUAAAUGUUUAAAGUUUUAUUCUAAAAUUAGUGAGUAUAGACAUUAUUGUUAUAGUUUGUAUGGAACAACAUUAGAACAGCAAUACAACAAACCGUUGUGUGAUUUGUUAAUUAGGUGUAUUAACUGUCAAAAGCCACUGUGUCCUGAAGAAAAGCAAAGACAUCUGGACAAAAAGCAAAGAUUCCAUAAUAUAAGGGGUCGGUGGACCGGUCGAUGUAUGUCUUGUUGCAGAUCAUCAAGAACACGUAGAGAAACCCAGCUGUAA318- example_title: NRAS proto-oncogene319  mask_index: 36320  mask_index_1based: 37321  masked_char: A322  output:323  - label: C324    score: 0.249393325  - label: Y326    score: 0.094936327  - label: S328    score: 0.059246329  - label: M330    score: 0.056208331  - label: B332    score: 0.050245333  pipeline_tag: fill-mask334  sequence_type: 5' UTR335  task: fill-mask336  text: GGGGCCGGAAGUGCCGCUCCUUGGUGGGGGCUGUUC<mask>UGGCGGUUCCGGGGUCUCCAACAUUUUUCCCGGCUGUGGUCCUAAAUCUGUCCAAAGCAGAGGCAGUGGAGCUUGAGGUUCUUGCUGGUGUGAA337- example_title: amyloid beta precursor protein338  mask_index: 15339  mask_index_1based: 16340  masked_char: A341  output:342  - label: G343    score: 0.092117344  - label: S345    score: 0.066339346  - label: R347    score: 0.05872348  - label: V349    score: 0.054818350  - label: K351    score: 0.04943352  pipeline_tag: fill-mask353  sequence_type: 5' UTR354  task: fill-mask355  text: GUCAGUUUCCUCGGC<mask>GCGGUAGGCGAGAGCACGCGGAGGAGCGUGCGCGGGGGCCCCGGGAGACGGCGGCGGUGGCGGCGCGGGCAGAGCAAGGACGCGGCGGAUCCCACUCGCACAGCAGCGCACUCGGUGCCCCGCGCAGGGUCGCG356- example_title: RUNX family transcription factor 1357  mask_index: 15358  mask_index_1based: 16359  masked_char: A360  output:361  - label: A362    score: 0.067454363  - label: W364    score: 0.064112365  - label: M366    score: 0.063179367  - label: H368    score: 0.062422369  - label: U370    score: 0.060935371  pipeline_tag: fill-mask372  sequence_type: 5' UTR373  task: fill-mask374  text: ACUUCUUUGGGCCUC<mask>UAAACAACCACAGAACCACAAGUUGGGUAGCCUGGCAGUGUCAGAAGUCUGAACCCAGCAUAGUGGUCAGCAGGCAGGACGAAUCACACUGAAUGCAAACCACAGGGUUUCGCAGCGUGGUAAAAGAAAUCAUUGAGUCCCCCGCCUUCAGAAGAGGGUGCAUUUUCAGGAGGAAGCG375- example_title: fragile X messenger ribonucleoprotein 1376  mask_index: 15377  mask_index_1based: 16378  masked_char: A379  output:380  - label: G381    score: 0.08842382  - label: S383    score: 0.071981384  - label: C385    score: 0.058599386  - label: V387    score: 0.057824388  - label: R389    score: 0.05744390  pipeline_tag: fill-mask391  sequence_type: 5' UTR392  task: fill-mask393  text: CUCAGUCAGGCGCUC<mask>GCUCCGUUUCGGUUUCACUUCCGGUGGAGGGCCGCCUCUGAGCGGGCGGCGGGCCGACGGCGAGCGCGGGCGGCGGCGGUGACGGAGGCGCCGCUGCCAGGGGGCGUGCGGCAGCGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGAGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCUGGGCCUCGAGCGCCCGCAGCCCACCUCUCGGGGGCGGGCUCCCGGCGCUAGCAGGGCUGAAGAGAAG394- example_title: MYC proto-oncogene395  mask_index: 10396  mask_index_1based: 11397  masked_char: A398  output:399  - label: U400    score: 0.06839401  - label: K402    score: 0.066684403  - label: G404    score: 0.065021405  - label: D406    score: 0.054979407  - label: W408    score: 0.050555409  pipeline_tag: fill-mask410  sequence_type: 5' UTR411  task: fill-mask412  text: AACUCGCUGU<mask>GUAAUUCCAGCGAGAGGCAGAGGGAGCGAGCGGGCGGCCGGCUAGGGUGGAAGAGCCGGGCGAGCAGAGCUGCGCUGCGGGCGUCCUGGGAAGGGAGAUCCGGAGCGAAUAGGGGGCUUCGCCUCUGGCCCAGCCCUCCCGCUGAUCCCCCAGCCAGCGGUCCGCAACCCUUGCCGCAUCCACGAAACUUUGCCCAUAGCAGCGGGCGGGCACUUUGCACUGGAACUUACAACACCCGAGCAAGGACGCGACUCUCCCGACGCGGGGAGGCUAUUCUGCCCAUUUGGGGACACUUCCCCGCCGCUGCCAGGACCCGCUUCUCUGAAAGGCUCUCCUUGCAGCUGCUUAGACG413- example_title: activating transcription factor 4414  mask_index: 20415  mask_index_1based: 21416  masked_char: A417  output:418  - label: C419    score: 0.083975420  - label: Y421    score: 0.068627422  - label: S423    score: 0.059391424  - label: B425    score: 0.058267426  - label: U427    score: 0.056083428  pipeline_tag: fill-mask429  sequence_type: 5' UTR430  task: fill-mask431  text: CAUUUCUACUUUGCCCGCCC<mask>CAGAUGUAGUUUUCUCUGCGCGUGUGCGUUUUCCCUCCUCCCCGCCCUCAGGGUCCACGGCCACCAUGGCGUAUUAGGGGCAGCAGUGCCUGCGGCAGCAUUGGCCUUUGCAGCGGCGGCAGCAGCACCAGGCUCUGCAGCGGCAACCCCCAGCGGCUUAAGCCAUGGCGCUUCUCACGGCAUUCAGCAGCAGCGUUGCUGUAACCGACAAAGACACCUUCGAAUUAAGCACAUUCCUCGAUUCCAGCAAAGCACCGCAAC432- example_title: Human GPI protein p137433  mask_index: 11434  mask_index_1based: 12435  masked_char: A436  output:437  - label: A438    score: 0.084748439  - label: M440    score: 0.061405441  - label: R442    score: 0.056181443  - label: W444    score: 0.052638445  - label: V446    score: 0.051978447  pipeline_tag: fill-mask448  sequence_type: 3' UTR449  task: fill-mask450  text: UUUUUAAAAGG<mask>AAAGAUACCAAAUGCCUGCUGCUACCACCCUUUUCAAUUGCUAUGUUUUGAAAGGCACCAGUAUGUGUUUUAGAUUGAUUUAAAUGUUUCAUUUAAAUCACGGACAGUAGUUUCAGUUCUGAUGGUAUAAGCAAAACAAAUAAAACGUUUAUAAAAGUUGUAUCUUGAAACACUGGUGUUCAACAGCUAGCAGCUUAUGUGAUUCACCCCAUGCCACGUUAGUGUCACAAAUUUUAUGGUUUAUCUCCAGCAACAUUUCUCUAGUACUUGCACUUAUUAUCUGAAUUC451- example_title: nucleophosmin 1452  mask_index: 11453  mask_index_1based: 12454  masked_char: A455  output:456  - label: U457    score: 0.07761458  - label: W459    score: 0.065915460  - label: A461    score: 0.055983462  - label: D463    score: 0.053178464  - label: K465    score: 0.051828466  pipeline_tag: fill-mask467  sequence_type: 3' UTR468  task: fill-mask469  text: GAAAAUAGUUU<mask>AACAAUUUGUUAAAAAAUUUUCCGUCUUAUUUCAUUUCUGUAACAGUUGAUAUCUGGCUGUCCUUUUUAUAAUGCAGAGUGAGAACUUUCCCUACCGUGUUUGAUAAAUGUUGUCCAGGUUCUAUUGCCAAGAAUGUGUUGUCCAAAAUGCCUGUUUAGUUUUUAAAGAUGGAACUCCACCCUUUGCUUGGUUUUAAGUAUGUAUGGAAUGUUAUGAUAGGACAUAGUAGUAGCGGUGGUCAGACAUGGAAAUGGUGGGGAGACAAAAAUAUACAUGUGAAAUAAAACUCAGUAUUUUAAUAAAGUAGCACGGUUUCUAUUGA470- example_title: superoxide dismutase 1471  mask_index: 12472  mask_index_1based: 13473  masked_char: A474  output:475  - label: U476    score: 0.062892477  - label: Y478    score: 0.057641479  - label: W480    score: 0.05284481  - label: H482    score: 0.052837483  - label: C484    score: 0.052829485  pipeline_tag: fill-mask486  sequence_type: 3' UTR487  task: fill-mask488  text: ACAUUCCCUUGG<mask>UGUAGUCUGAGGCCCCUUAACUCAUCUGUUAUCCUGCUAGCUGUAGAAAUGUAUCCUGAUAAACAUUAAACACUGUAAUCUUAAAAGUGUAAUUGUGUGACUUUUUCAGAGUUGCUUUAAAGUACCUGUAGUGAGAAACUGAUUUAUGAUCACUUGGAAGAUUUGUAUAGUUUUAUAAAACUCAGUUAAAAUGUCUGUUUCAAUGACCUGUAUUUUGCCAGACUUAAAUCACAGAUGGGUAUUAAACUUGUCAGAAUUUCUUUGUCAUUCAAGCCUGUGAAUAAAAACCCUGUAUGGCACUUAUUAUGAGGCUAUUAAAAGAAUCCAAAUUCAAACUAAA489- example_title: hemoglobin subunit alpha 2490  mask_index: 13491  mask_index_1based: 14492  masked_char: A493  output:494  - label: G495    score: 0.32064496  - label: R497    score: 0.073522498  - label: S499    score: 0.068353500  - label: K501    score: 0.064896502  - label: V503    score: 0.042866504  pipeline_tag: fill-mask505  sequence_type: 3' UTR506  task: fill-mask507  text: CUGGAGCCUCGGU<mask>GCCGUUCCUCCUGCCCGCUGGGCCUCCCAACGGGCCCUCCUCCCCUCCUUGCACCGGCCCUUCCUGGUCUUUGAAUAAAGUCUGAGUGGGCAGCA508- example_title: BRAF proto-oncogene509  mask_index: 12510  mask_index_1based: 13511  masked_char: A512  output:513  - label: A514    score: 0.130968515  - label: R516    score: 0.091002517  - label: W518    score: 0.085338519  - label: D520    score: 0.077222521  - label: G522    score: 0.063232523  pipeline_tag: fill-mask524  sequence_type: 3' UTR525  task: fill-mask526  text: AACAAAUGAGUG<mask>GAGAGUUCAGGAGAGUAGCAACAAAAGGAAAAUAAAUGAACAUAUGUUUGCUUAUAUGUUAAAUUGAAUAAAAUACUCUCUUUUUUUUUAAGGUGAACCAAAGAACACUUGUGUGGUUAAAGACUAGAUAUAAUUUUUCCCCAAACUAAAAUUUAUACUUAACAUUGGAUUUUUAACAUCCAAGGGUUAAAAUACAUAGACAUUGCUAAAAAUUGGCAGAGCCUCUUCUAGAGGCUUUACUUUCUGUUCCGGGUUUGUAUCAUUCACUUGGUUAUUUUAAGUAGUAAACUUCAGUUUCUCAUGCAACUUUUGUUGCCAGCUAUCACAUGUCCACUAGGGACUCCAGAAGAAGACCCUACCUAUGCCUGUGUUUGCAGGUGAGAAGUUGGCAGUCGGUUAGCCUGGG527- example_title: H3 clustered histone 1528  mask_index: 17529  mask_index_1based: 18530  masked_char: A531  output:532  - label: A533    score: 0.06926534  - label: R535    score: 0.05768536  - label: W537    score: 0.052921538  - label: D539    score: 0.05124540  - label: M541    score: 0.048795542  pipeline_tag: fill-mask543  sequence_type: 3' UTR544  task: fill-mask545  text: UUACUGUGGUCUCUCUG<mask>CGGUCCAAGCAAAGGCUCUUUUCAGAGCCACCACCUUUUC546---547 548# SpliceBERT549 550Pre-trained model on messenger RNA precursor (pre-mRNA) using a masked language modeling (MLM) objective.551 552## Disclaimer553 554This is an UNOFFICIAL implementation of the [Self-supervised learning on millions of pre-mRNA sequences improves sequence-based RNA splicing prediction](https://doi.org/10.1101/2023.01.31.526427) by Ken Chen, et al.555 556The OFFICIAL repository of SpliceBERT is at [chenkenbio/SpliceBERT](https://github.com/chenkenbio/SpliceBERT).557 558> [!TIP]559> The MultiMolecule team has confirmed that the provided model and checkpoints are producing the same intermediate representations as the original implementation.560 561**The team releasing SpliceBERT did not write this model card for this model so this model card has been written by the MultiMolecule team.**562 563## Model Details564 565SpliceBERT is a [bert](https://huggingface.co/google-bert/bert-base-uncased)-style model pre-trained on a large corpus of messenger RNA precursor sequences in a self-supervised fashion. This means that the model was trained on the raw nucleotides of RNA sequences only, with an automatic process to generate inputs and labels from those texts. Please refer to the [Training Details](#training-details) section for more information on the training process.566 567### Variants568 569- **[multimolecule/splicebert](https://huggingface.co/multimolecule/splicebert)**: The SpliceBERT model.570- **[multimolecule/splicebert.510](https://huggingface.co/multimolecule/splicebert.510)**: The intermediate SpliceBERT model.571- **[multimolecule/splicebert-human.510](https://huggingface.co/multimolecule/splicebert-human.510)**: The intermediate SpliceBERT model pre-trained on human data only.572 573### Model Specification574 575<table>576<thead>577  <tr>578    <th>Variants</th>579    <th>Num Layers</th>580    <th>Hidden Size</th>581    <th>Num Heads</th>582    <th>Intermediate Size</th>583    <th>Num Parameters (M)</th>584    <th>FLOPs (G)</th>585    <th>MACs (G)</th>586    <th>Max Num Tokens</th>587  </tr>588</thead>589<tbody>590  <tr>591    <td>splicebert</td>592    <td rowspan="3">6</td>593    <td rowspan="3">512</td>594    <td rowspan="3">16</td>595    <td rowspan="3">2048</td>596    <td>19.72</td>597    <td>22.66</td>598    <td>11.27</td>599    <td>1024</td>600  </tr>601  <tr>602    <td><b>splicebert.510</b></td>603    <td rowspan="2">19.45</td>604    <td rowspan="2">22.56</td>605    <td rowspan="2">11.22</td>606    <td rowspan="2">510</td>607  </tr>608  <tr>609    <td>splicebert-human.510</td>610  </tr>611</tbody>612</table>613 614### Links615 616- **Code**: [multimolecule.splicebert](https://github.com/DLS5-Omics/multimolecule/tree/master/multimolecule/models/splicebert)617- **Data**: [UCSC Genome Browser](https://genome.ucsc.edu)618- **Paper**: [Self-supervised learning on millions of pre-mRNA sequences improves sequence-based RNA splicing prediction](https://doi.org/10.1101/2023.01.31.526427)619- **Developed by**: Ken Chen, Yue Zhou, Maolin Ding, Yu Wang, Zhixiang Ren, Yuedong Yang620- **Model type**: [BERT](https://huggingface.co/google-bert/bert-base-uncased)621- **Original Repository**: [chenkenbio/SpliceBERT](https://github.com/chenkenbio/SpliceBERT)622 623## Usage624 625The model file depends on the [`multimolecule`](https://multimolecule.danling.org) library. You can install it using pip:626 627```bash628pip install multimolecule629```630 631### Direct Use632 633#### Masked Language Modeling634 635You can use this model directly with a pipeline for masked language modeling:636 637```python638import multimolecule  # you must import multimolecule to register models639from transformers import pipeline640 641predictor = pipeline("fill-mask", model="multimolecule/splicebert")642output = predictor("gguc<mask>cucugguuagaccagaucugagccu")643```644 645### Downstream Use646 647#### Extract Features648 649Here is how to use this model to get the features of a given sequence in PyTorch:650 651```python652from multimolecule import RnaTokenizer, SpliceBertModel653 654 655tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")656model = SpliceBertModel.from_pretrained("multimolecule/splicebert")657 658text = "UAGCUUAUCAGACUGAUGUUG"659input = tokenizer(text, return_tensors="pt")660 661output = model(**input)662```663 664#### Sequence Classification / Regression665 666> [!NOTE]667> This model is not fine-tuned for any specific task. You will need to fine-tune the model on a downstream task to use it for sequence classification or regression.668 669Here is how to use this model as backbone to fine-tune for a sequence-level task in PyTorch:670 671```python672import torch673from multimolecule import RnaTokenizer, SpliceBertForSequencePrediction674 675 676tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")677model = SpliceBertForSequencePrediction.from_pretrained("multimolecule/splicebert")678 679text = "UAGCUUAUCAGACUGAUGUUG"680input = tokenizer(text, return_tensors="pt")681label = torch.tensor([1])682 683output = model(**input, labels=label)684```685 686#### Token Classification / Regression687 688> [!NOTE]689> This model is not fine-tuned for any specific task. You will need to fine-tune the model on a downstream task to use it for token classification or regression.690 691Here is how to use this model as backbone to fine-tune for a nucleotide-level task in PyTorch:692 693```python694import torch695from multimolecule import RnaTokenizer, SpliceBertForTokenPrediction696 697 698tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")699model = SpliceBertForTokenPrediction.from_pretrained("multimolecule/splicebert")700 701text = "UAGCUUAUCAGACUGAUGUUG"702input = tokenizer(text, return_tensors="pt")703label = torch.randint(2, (len(text), ))704 705output = model(**input, labels=label)706```707 708#### Contact Classification / Regression709 710> [!NOTE]711> This model is not fine-tuned for any specific task. You will need to fine-tune the model on a downstream task to use it for contact classification or regression.712 713Here is how to use this model as backbone to fine-tune for a contact-level task in PyTorch:714 715```python716import torch717from multimolecule import RnaTokenizer, SpliceBertForContactPrediction718 719 720tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")721model = SpliceBertForContactPrediction.from_pretrained("multimolecule/splicebert")722 723text = "UAGCUUAUCAGACUGAUGUUG"724input = tokenizer(text, return_tensors="pt")725label = torch.randint(2, (len(text), len(text)))726 727output = model(**input, labels=label)728```729 730## Training Details731 732SpliceBERT used Masked Language Modeling (MLM) as the pre-training objective: taking a sequence, the model randomly masks 15% of the tokens in the input then runs the entire masked sentence through the model and has to predict the masked tokens. This is comparable to the Cloze task in language modeling.733 734### Training Data735 736The SpliceBERT model was pre-trained on messenger RNA precursor sequences from [UCSC Genome Browser](https://genome.ucsc.edu).737UCSC Genome Browser provides visualization, analysis, and download of comprehensive vertebrate genome data with aligned annotation tracks (known genes, predicted genes, ESTs, mRNAs, CpG islands, etc.).738 739SpliceBERT collected reference genomes and gene annotations from the UCSC Genome Browser for 72 vertebrate species. It applied [bedtools getfasta](https://bedtools.readthedocs.io/en/latest/content/tools/getfasta.html) to extract pre-mRNA sequences from the reference genomes based on the gene annotations. The pre-mRNA sequences are then used to pre-train SpliceBERT. The pre-training data contains 2 million pre-mRNA sequences with a total length of 65 billion nucleotides.740 741Note [`RnaTokenizer`][multimolecule.RnaTokenizer] will convert "T"s to "U"s for you, you may disable this behaviour by passing `replace_T_with_U=False`.742 743### Training Procedure744 745#### Preprocessing746 747SpliceBERT used masked language modeling (MLM) as the pre-training objective. The masking procedure is similar to the one used in BERT:748 749- Mask rate: 15%750- Replacement: `<mask>` for 80% of masked tokens751- Replacement: random token for 10% of masked tokens752- Replacement: unchanged token for 10% of masked tokens753 754#### Pre-training755 756The model was trained on 8 NVIDIA V100 GPUs.757 758- Optimizer: AdamW759- Learning rate: 1e-4760- Learning rate scheduler: ReduceLROnPlateau(patience=3)761 762SpliceBERT trained model in a two-stage training process:763 7641. Pre-train with sequences of a fixed length of 510 nucleotides.7652. Pre-train with sequences of a variable length between 64 and 1024 nucleotides.766 767The intermediate model after the first stage is available as `multimolecule/splicebert.510`.768 769SpliceBERT also pre-trained a model on human data only to validate the contribution of multi-species pre-training. The intermediate model after the first stage is available as `multimolecule/splicebert-human.510`.770 771## Citation772 773```bibtex774@article {chen2023self,775	author = {Chen, Ken and Zhou, Yue and Ding, Maolin and Wang, Yu and Ren, Zhixiang and Yang, Yuedong},776	title = {Self-supervised learning on millions of pre-mRNA sequences improves sequence-based RNA splicing prediction},777	elocation-id = {2023.01.31.526427},778	year = {2023},779	doi = {10.1101/2023.01.31.526427},780	publisher = {Cold Spring Harbor Laboratory},781	abstract = {RNA splicing is an important post-transcriptional process of gene expression in eukaryotic cells. Predicting RNA splicing from primary sequences can facilitate the interpretation of genomic variants. In this study, we developed a novel self-supervised pre-trained language model, SpliceBERT, to improve sequence-based RNA splicing prediction. Pre-training on pre-mRNA sequences from vertebrates enables SpliceBERT to capture evolutionary conservation information and characterize the unique property of splice sites. SpliceBERT also improves zero-shot prediction of variant effects on splicing by considering sequence context information, and achieves superior performance for predicting branchpoint in the human genome and splice sites across species. Our study highlighted the importance of pre-training genomic language models on a diverse range of species and suggested that pre-trained language models were promising for deciphering the sequence logic of RNA splicing.Competing Interest StatementThe authors have declared no competing interest.},782	URL = {https://www.biorxiv.org/content/early/2023/05/09/2023.01.31.526427},783	eprint = {https://www.biorxiv.org/content/early/2023/05/09/2023.01.31.526427.full.pdf},784	journal = {bioRxiv}785}786```787 788> [!NOTE]789> The artifacts distributed in this repository are part of the MultiMolecule project.790> If MultiMolecule supports your research, please cite the MultiMolecule project as follows:791 792```bibtex793@software{chen_2024_12638419,794  author    = {Chen, Zhiyuan and Zhu, Sophia Y.},795  title     = {MultiMolecule},796  doi       = {10.5281/zenodo.12638419},797  publisher = {Zenodo},798  url       = {https://doi.org/10.5281/zenodo.12638419},799  year      = 2024,800  month     = may,801  day       = 4802}803```804 805## Contact806 807Please use GitHub issues of [MultiMolecule](https://github.com/DLS5-Omics/multimolecule/issues) for any questions or comments on the model card.808 809Please contact the authors of the [SpliceBERT paper](https://doi.org/10.1101/2023.01.31.526427) for questions or comments on the paper/model.810 811## License812 813This model implementation is licensed under the [GNU Affero General Public License](license.md).814 815For additional terms and clarifications, please refer to our [License FAQ](license-faq.md).816 817```spdx818SPDX-License-Identifier: AGPL-3.0-or-later819```