multimolecule/splicebert.510
067
1---2datasets:3- multimolecule/ucsc-genome-browser4library_name: multimolecule5license: agpl-3.06mask_token: <mask>7pipeline_tag: fill-mask8tags:9- Biology10- RNA11- ncRNA12- rna13widget:14- example_title: microRNA 2115 mask_index: 1116 mask_index_1based: 1217 masked_char: A18 output:19 - label: C20 score: 0.07627921 - label: S22 score: 0.06061423 - label: M24 score: 0.05649125 - label: V26 score: 0.05356727 - label: Y28 score: 0.04820329 pipeline_tag: fill-mask30 sequence_type: ncRNA31 task: fill-mask32 text: UAGCUUAUCAG<mask>CUGAUGUUGA33- example_title: microRNA 146a34 mask_index: 1035 mask_index_1based: 1136 masked_char: A37 output:38 - label: G39 score: 0.16400540 - label: R41 score: 0.08606942 - label: K43 score: 0.06784844 - label: D45 score: 0.05924346 - label: A47 score: 0.04516948 pipeline_tag: fill-mask49 sequence_type: ncRNA50 task: fill-mask51 text: UGAGAACUGA<mask>UUCCAUGGGUU52- example_title: microRNA 15553 mask_index: 1554 mask_index_1based: 1655 masked_char: A56 output:57 - label: U58 score: 0.06264659 - label: W60 score: 0.05326461 - label: K62 score: 0.05004963 - label: Y64 score: 0.04972265 - label: D66 score: 0.04840867 pipeline_tag: fill-mask68 sequence_type: ncRNA69 task: fill-mask70 text: UUAAUGCUAAUCGUG<mask>UAGGGGUU71- example_title: RNA component of mitochondrial RNA processing endoribonuclease72 mask_index: 1173 mask_index_1based: 1274 masked_char: A75 output:76 - label: C77 score: 0.07944578 - label: S79 score: 0.06397580 - label: M81 score: 0.05370982 - label: V83 score: 0.05296984 - label: G85 score: 0.05151886 pipeline_tag: fill-mask87 sequence_type: ncRNA88 task: fill-mask89 text: GGUUCGUGCUG<mask>AGGCCUGUAUCCUAGGCUACACACUGAGGACUCUGUUCCUCCCCUUUCCGCCUAGGGGAAAGUCCCCGGACCUCGGGCAGAGAGUGCCACGUGCAUACGCACGUAGACAUUCCCCGCUUCCCACUCCAAAGUCCGCCAAGAAGCGUAUCCCGCUGAGCGGCGUGGCGCGGGGGCGUCAUCCGUCAGCUCCCUCUAGUUACGCAGGCAGUGCGUGUCCGCGCACCAACCACACGGGGCUCAUUCUCAGCGCGGCUGUAAAAAAAAA90- example_title: 7SK small nuclear RNA91 mask_index: 1392 mask_index_1based: 1493 masked_char: A94 output:95 - label: A96 score: 0.06808697 - label: R98 score: 0.06639399 - label: G100 score: 0.064741101 - label: V102 score: 0.057095103 - label: M104 score: 0.053618105 pipeline_tag: fill-mask106 sequence_type: ncRNA107 task: fill-mask108 text: GGAUGUGAGGGCG<mask>UCUGGCUGCGACAUCUGUCACCCCAUUGAUCGCCAGGGUUGAUUCGGCUGAUCUGGCUGGCUAGGCGGGUGUCCCCUUCCUCCCUCACCGCUCCAUGUGCGUCCCUCCCGAAGCUGCGCGCUCGGUCGAAGAGGACGACCAUCCCCGAUAGAGGAGGACCGGUCUUCGGUCAAGGGUAUACGAGUAGCUGCGCUCCCCUGCUAGAACCUCCAAACAAGCUCUCAAGGUCCAUUUGUAGGAGAACGUAGGGUAGUCAAGCUUCCAAGACUCCAGACACAUCCAAAUGAGGCGCUGCAUGUGGCAGUCUGCCUUUCUUUU109- example_title: telomerase RNA component110 mask_index: 23111 mask_index_1based: 24112 masked_char: A113 output:114 - label: C115 score: 0.061932116 - label: Y117 score: 0.061489118 - label: U119 score: 0.061049120 - label: H121 score: 0.060014122 - label: M123 score: 0.059503124 pipeline_tag: fill-mask125 sequence_type: ncRNA126 task: fill-mask127 text: GGGUUGCGGAGGGUGGGCCUGGG<mask>GGGGUGGUGGCCAUUUUUUGUCUAACCCUAACUGAGAAGGGCGUAGGCGCCGUGCUUUUGCUCCCCGCGCGCUGUUUUUCUCGCUGACUUUCAGCGGGCGGAAAAGCCUCGGCCUGCCGCCUUCCACCGUUCAUUCUAGAGCAAACAAAAAAUGUCAGCUGCUGGCCCGUUCGCCCCUCCCGGGGACCUGCGGCGGGUCGCCUGCCCAGCCCCCGAACCCCGCCUGGAGGCCGCGGUCGGCCCGGGGCUUCUCCGGAGGCACCCACUGCCACCGCGAAGAGUUGGGCUCUGUCAGCCGCGGGUCUCUCGGGGGCGAGGGCGAGGUUCAGGCCUUUCAGGCCGCAGGAAGAGGAACGGAGCGAGUCCCCGCGCGCGGCGCGAUUCCCUGAGCUGUGGGACGUGCACCCAGGACUCGGCUCACACAUGC128- example_title: vault RNA 2-1129 mask_index: 12130 mask_index_1based: 13131 masked_char: A132 output:133 - label: U134 score: 0.074338135 - label: K136 score: 0.062814137 - label: Y138 score: 0.053943139 - label: B140 score: 0.053653141 - label: G142 score: 0.053077143 pipeline_tag: fill-mask144 sequence_type: ncRNA145 task: fill-mask146 text: CGGGUCGGAGUU<mask>GCUCAAGCGGUUACCUCCUCAUGCCGGACUUUCUAUCUGUCCAUCUCUGUGCUGGGGUUCGAGACCCGCGGGUGCUUACUGACCCUUUUAUGCAA147- example_title: brain cytoplasmic RNA 1148 mask_index: 18149 mask_index_1based: 19150 masked_char: A151 output:152 - label: A153 score: 0.416301154 - label: R155 score: 0.08191156 - label: M157 score: 0.060285158 - label: W159 score: 0.052732160 - label: I161 score: 0.039218162 pipeline_tag: fill-mask163 sequence_type: ncRNA164 task: fill-mask165 text: GGCCGGGCGCGGUGGCUC<mask>CGCCUGUAAUCCCAGCUCUCAGGGAGGCUAAGAGGCGGGAGGAUAGCUUGAGCCCAGGAGUUCGAGACCUGCCUGGGCAAUAUAGCGAGACCCCGUUCUCCAGAAAAAGGAAAAAAAAAAACAAAAGACAAAAAAAAAAUAAGCGUAACUUCCCUCAAAGCAACAACCCCCCCCCCCCUUU166- example_title: HIV-1 TAR-WT167 mask_index: 13168 mask_index_1based: 14169 masked_char: A170 output:171 - label: A172 score: 0.087018173 - label: R174 score: 0.078768175 - label: W176 score: 0.075575177 - label: D178 score: 0.074122179 - label: G180 score: 0.0713181 pipeline_tag: fill-mask182 sequence_type: ncRNA183 task: fill-mask184 text: GGUCUCUCUGGUU<mask>GACCAGAUCUGAGCCUGGGAGCUCUCUGGCUAACUAGGGAACC185- example_title: prion protein (Kanno blood group)186 mask_index: 21187 mask_index_1based: 22188 masked_char: A189 output:190 - label: C191 score: 0.203007192 - label: S193 score: 0.088737194 - label: Y195 score: 0.059722196 - label: M197 score: 0.0563198 - label: B199 score: 0.051719200 pipeline_tag: fill-mask201 sequence_type: mRNA202 task: fill-mask203 text: AUGGCGAACCUUGGCUGCUGG<mask>UGCUGGUUCUCUUUGUGGCCACAUGGAGUGACCUGGGCCUCUGC204- example_title: interleukin 10205 mask_index: 11206 mask_index_1based: 12207 masked_char: A208 output:209 - label: U210 score: 0.145719211 - label: W212 score: 0.092637213 - label: A214 score: 0.058892215 - label: H216 score: 0.054023217 - label: Y218 score: 0.051742219 pipeline_tag: fill-mask220 sequence_type: mRNA221 task: fill-mask222 text: AUGCACAGCUC<mask>GCACUGCUCUGUUGCCUGGUCCUCCUGACUGGGGUGAGGGCC223- example_title: Zaire ebolavirus224 mask_index: 11225 mask_index_1based: 12226 masked_char: A227 output:228 - label: U229 score: 0.089489230 - label: W231 score: 0.088384232 - label: A233 score: 0.087294234 - label: H235 score: 0.065665236 - label: Y237 score: 0.056951238 pipeline_tag: fill-mask239 sequence_type: mRNA240 task: fill-mask241 text: AAUGUUCAAAC<mask>CUUUGUGAAGCUCUGUUAGCUGAUGGUCUUGCUAAAGCAUUUCCUAGCAAUAUGAUGGUAGUCACAGAGCGUGAGCAAAAAGAAAGCUUAUUGCAUCAAGCAUCAUGGCACCACACAAGUGAUGAUUUUGGUGAGCAUGCCACAGUUAGAGGGAGUAGCUUUGUAACUGAUUUAGAGAAAUACAAUCUUGCAUUUAGAUAUGAGUUUACAGCACCUUUUAUAGAAUAUUGUAACCGUUGCUAUGGUGUUAAGAAUGUUUUUAAUUGGAUGCAUUAUACAAUCCCACAGUGUUAU242- example_title: SARS coronavirus243 mask_index: 14244 mask_index_1based: 15245 masked_char: A246 output:247 - label: U248 score: 0.124725249 - label: Y250 score: 0.065713251 - label: W252 score: 0.065298253 - label: K254 score: 0.057034255 - label: H256 score: 0.05285257 pipeline_tag: fill-mask258 sequence_type: mRNA259 task: fill-mask260 text: AUGUUUAUUUUCUU<mask>UUAUUUCUUACUCUCACUAGUGGUAGUGACCUUGACCGGUGCACCACUUUUGAUGAUGUUCAAGCUCCUAAUUACACUCAACAUACUUCAUCUAUGAGGGGGGUUUACUAUCCUGAUGAAAUUUUUAGAUCAGACACUCUUUAUUUAACUCAGGAUUUAUUUCUUCCAUUUUAUUCUAAUGUUACAGGGUUUCAUACUAUUAAUCAUACGUUUGACAACCCUGUCAUACCUUUUAAGGAUGGUAUUUAUUUUGCUGCCACAGAGAAAUCAAAUGUUGUCCGUGGUUGGGUUUUUGGUUCUACCAUGAACAACAAGUCACAGUCGGUGAUUAUUAUUAACAAUUCUACUAAUGUUGUUAUACGAGCAUGUAACUUUGAAUUGUGUGACAACCCUUUCUUUGCUGUUUCUAAACCCAUGGGUACACAGACACAUACUAUGAUAUUCGAUAAUGCAUUUAAAUGCACUUUCGAGUACAUAUCU261- example_title: insulin262 mask_index: 12263 mask_index_1based: 13264 masked_char: A265 output:266 - label: C267 score: 0.280004268 - label: S269 score: 0.107256270 - label: M271 score: 0.057172272 - label: Y273 score: 0.053979274 - label: V275 score: 0.05121276 pipeline_tag: fill-mask277 sequence_type: mRNA278 task: fill-mask279 text: AUGGCCCUGUGG<mask>UGCGCCUCCUGCCCCUGCUGGCGCUGCUGGCCCUCUGGGGACCUGACCCAGCCGCAGCCUUUGUGAACCAACACCUGUGCGGCUCACACCUGGUGGAAGCUCUCUACCUAGUGUGCGGGGAACGAGGCUUCUUCUACACACCCAAGACCCGCCGGGAGGCAGAGGACCUGCAGGUGGGGCAGGUGGAGCUGGGCGGGGGCCCUGGUGCAGGCAGCCUGCAGCCCUUGGCCCUGGAGGGGUCCCUGCAGAAGCGUGGCAUUGUGGAACAAUGCUGUACCAGCAUCUGCUCCCUCUACCAGCUGGAGAACUACUGCAACUAG280- example_title: cyclin dependent kinase inhibitor 2A281 mask_index: 18282 mask_index_1based: 19283 masked_char: A284 output:285 - label: G286 score: 0.098564287 - label: S288 score: 0.08414289 - label: C290 score: 0.071827291 - label: V292 score: 0.066422293 - label: R294 score: 0.063874295 pipeline_tag: fill-mask296 sequence_type: mRNA297 task: fill-mask298 text: AUGGAGCCGGCGGCGGGG<mask>GCAGCAUGGAGCCUUCGGCUGACUGGCUGGCCACGGCCGCGGCCCGGGGUCGGGUAGAGGAGGUGCGGGCGCUGCUGGAGGCGGGGGCGCUGCCCAACGCACCGAAUAGUUACGGUCGGAGGCCGAUCCAGGUCAUGAUGAUGGGCAGCGCCCGAGUGGCGGAGCUGCUGCUGCUCCACGGCGCGGAGCCCAACUGCGCCGACCCCGCCACUCUCACCCGACCCGUGCACGACGCUGCCCGGGAGGGCUUCCUGGACACGCUGGUGGUGCUGCACCGGGCCGGGGCGCGGCUGGACGUGCGCGAUGCCUGGGGCCGUCUGCCCGUGGACCUGGCUGAGGAGCUGGGCCAUCGCGAUGUCGCACGGUACCUGCGCGCGGCUGCGGGGGGCACCAGAGGCAGUAACCAUGCCCGCAUAGAUGCCGCGGAAGGUCCCUCAGACAUCCCCGAUUGA299- example_title: human papillomavirus type 16 E6300 mask_index: 10301 mask_index_1based: 11302 masked_char: A303 output:304 - label: A305 score: 0.086406306 - label: W307 score: 0.077731308 - label: U309 score: 0.069927310 - label: D311 score: 0.068029312 - label: R313 score: 0.0671314 pipeline_tag: fill-mask315 sequence_type: mRNA316 task: fill-mask317 text: AUGCACCAAA<mask>GAGAACUGCAAUGUUUCAGGACCCACAGGAGCGACCCAGAAAGUUACCACAGUUAUGCACAGAGCUGCAAACAACUAUACAUGAUAUAAUAUUAGAAUGUGUGUACUGCAAGCAACAGUUACUGCGACGUGAGGUAUAUGACUUUGCUUUUCGGGAUUUAUGCAUAGUAUAUAGAGAUGGGAAUCCAUAUGCUGUAUGUGAUAAAUGUUUAAAGUUUUAUUCUAAAAUUAGUGAGUAUAGACAUUAUUGUUAUAGUUUGUAUGGAACAACAUUAGAACAGCAAUACAACAAACCGUUGUGUGAUUUGUUAAUUAGGUGUAUUAACUGUCAAAAGCCACUGUGUCCUGAAGAAAAGCAAAGACAUCUGGACAAAAAGCAAAGAUUCCAUAAUAUAAGGGGUCGGUGGACCGGUCGAUGUAUGUCUUGUUGCAGAUCAUCAAGAACACGUAGAGAAACCCAGCUGUAA318- example_title: NRAS proto-oncogene319 mask_index: 36320 mask_index_1based: 37321 masked_char: A322 output:323 - label: C324 score: 0.249393325 - label: Y326 score: 0.094936327 - label: S328 score: 0.059246329 - label: M330 score: 0.056208331 - label: B332 score: 0.050245333 pipeline_tag: fill-mask334 sequence_type: 5' UTR335 task: fill-mask336 text: GGGGCCGGAAGUGCCGCUCCUUGGUGGGGGCUGUUC<mask>UGGCGGUUCCGGGGUCUCCAACAUUUUUCCCGGCUGUGGUCCUAAAUCUGUCCAAAGCAGAGGCAGUGGAGCUUGAGGUUCUUGCUGGUGUGAA337- example_title: amyloid beta precursor protein338 mask_index: 15339 mask_index_1based: 16340 masked_char: A341 output:342 - label: G343 score: 0.092117344 - label: S345 score: 0.066339346 - label: R347 score: 0.05872348 - label: V349 score: 0.054818350 - label: K351 score: 0.04943352 pipeline_tag: fill-mask353 sequence_type: 5' UTR354 task: fill-mask355 text: GUCAGUUUCCUCGGC<mask>GCGGUAGGCGAGAGCACGCGGAGGAGCGUGCGCGGGGGCCCCGGGAGACGGCGGCGGUGGCGGCGCGGGCAGAGCAAGGACGCGGCGGAUCCCACUCGCACAGCAGCGCACUCGGUGCCCCGCGCAGGGUCGCG356- example_title: RUNX family transcription factor 1357 mask_index: 15358 mask_index_1based: 16359 masked_char: A360 output:361 - label: A362 score: 0.067454363 - label: W364 score: 0.064112365 - label: M366 score: 0.063179367 - label: H368 score: 0.062422369 - label: U370 score: 0.060935371 pipeline_tag: fill-mask372 sequence_type: 5' UTR373 task: fill-mask374 text: ACUUCUUUGGGCCUC<mask>UAAACAACCACAGAACCACAAGUUGGGUAGCCUGGCAGUGUCAGAAGUCUGAACCCAGCAUAGUGGUCAGCAGGCAGGACGAAUCACACUGAAUGCAAACCACAGGGUUUCGCAGCGUGGUAAAAGAAAUCAUUGAGUCCCCCGCCUUCAGAAGAGGGUGCAUUUUCAGGAGGAAGCG375- example_title: fragile X messenger ribonucleoprotein 1376 mask_index: 15377 mask_index_1based: 16378 masked_char: A379 output:380 - label: G381 score: 0.08842382 - label: S383 score: 0.071981384 - label: C385 score: 0.058599386 - label: V387 score: 0.057824388 - label: R389 score: 0.05744390 pipeline_tag: fill-mask391 sequence_type: 5' UTR392 task: fill-mask393 text: CUCAGUCAGGCGCUC<mask>GCUCCGUUUCGGUUUCACUUCCGGUGGAGGGCCGCCUCUGAGCGGGCGGCGGGCCGACGGCGAGCGCGGGCGGCGGCGGUGACGGAGGCGCCGCUGCCAGGGGGCGUGCGGCAGCGCGGCGGCGGCGGCGGCGGCGGCGGCGGCGGAGGCGGCGGCGGCGGCGGCGGCGGCGGCGGCUGGGCCUCGAGCGCCCGCAGCCCACCUCUCGGGGGCGGGCUCCCGGCGCUAGCAGGGCUGAAGAGAAG394- example_title: MYC proto-oncogene395 mask_index: 10396 mask_index_1based: 11397 masked_char: A398 output:399 - label: U400 score: 0.06839401 - label: K402 score: 0.066684403 - label: G404 score: 0.065021405 - label: D406 score: 0.054979407 - label: W408 score: 0.050555409 pipeline_tag: fill-mask410 sequence_type: 5' UTR411 task: fill-mask412 text: AACUCGCUGU<mask>GUAAUUCCAGCGAGAGGCAGAGGGAGCGAGCGGGCGGCCGGCUAGGGUGGAAGAGCCGGGCGAGCAGAGCUGCGCUGCGGGCGUCCUGGGAAGGGAGAUCCGGAGCGAAUAGGGGGCUUCGCCUCUGGCCCAGCCCUCCCGCUGAUCCCCCAGCCAGCGGUCCGCAACCCUUGCCGCAUCCACGAAACUUUGCCCAUAGCAGCGGGCGGGCACUUUGCACUGGAACUUACAACACCCGAGCAAGGACGCGACUCUCCCGACGCGGGGAGGCUAUUCUGCCCAUUUGGGGACACUUCCCCGCCGCUGCCAGGACCCGCUUCUCUGAAAGGCUCUCCUUGCAGCUGCUUAGACG413- example_title: activating transcription factor 4414 mask_index: 20415 mask_index_1based: 21416 masked_char: A417 output:418 - label: C419 score: 0.083975420 - label: Y421 score: 0.068627422 - label: S423 score: 0.059391424 - label: B425 score: 0.058267426 - label: U427 score: 0.056083428 pipeline_tag: fill-mask429 sequence_type: 5' UTR430 task: fill-mask431 text: CAUUUCUACUUUGCCCGCCC<mask>CAGAUGUAGUUUUCUCUGCGCGUGUGCGUUUUCCCUCCUCCCCGCCCUCAGGGUCCACGGCCACCAUGGCGUAUUAGGGGCAGCAGUGCCUGCGGCAGCAUUGGCCUUUGCAGCGGCGGCAGCAGCACCAGGCUCUGCAGCGGCAACCCCCAGCGGCUUAAGCCAUGGCGCUUCUCACGGCAUUCAGCAGCAGCGUUGCUGUAACCGACAAAGACACCUUCGAAUUAAGCACAUUCCUCGAUUCCAGCAAAGCACCGCAAC432- example_title: Human GPI protein p137433 mask_index: 11434 mask_index_1based: 12435 masked_char: A436 output:437 - label: A438 score: 0.084748439 - label: M440 score: 0.061405441 - label: R442 score: 0.056181443 - label: W444 score: 0.052638445 - label: V446 score: 0.051978447 pipeline_tag: fill-mask448 sequence_type: 3' UTR449 task: fill-mask450 text: UUUUUAAAAGG<mask>AAAGAUACCAAAUGCCUGCUGCUACCACCCUUUUCAAUUGCUAUGUUUUGAAAGGCACCAGUAUGUGUUUUAGAUUGAUUUAAAUGUUUCAUUUAAAUCACGGACAGUAGUUUCAGUUCUGAUGGUAUAAGCAAAACAAAUAAAACGUUUAUAAAAGUUGUAUCUUGAAACACUGGUGUUCAACAGCUAGCAGCUUAUGUGAUUCACCCCAUGCCACGUUAGUGUCACAAAUUUUAUGGUUUAUCUCCAGCAACAUUUCUCUAGUACUUGCACUUAUUAUCUGAAUUC451- example_title: nucleophosmin 1452 mask_index: 11453 mask_index_1based: 12454 masked_char: A455 output:456 - label: U457 score: 0.07761458 - label: W459 score: 0.065915460 - label: A461 score: 0.055983462 - label: D463 score: 0.053178464 - label: K465 score: 0.051828466 pipeline_tag: fill-mask467 sequence_type: 3' UTR468 task: fill-mask469 text: GAAAAUAGUUU<mask>AACAAUUUGUUAAAAAAUUUUCCGUCUUAUUUCAUUUCUGUAACAGUUGAUAUCUGGCUGUCCUUUUUAUAAUGCAGAGUGAGAACUUUCCCUACCGUGUUUGAUAAAUGUUGUCCAGGUUCUAUUGCCAAGAAUGUGUUGUCCAAAAUGCCUGUUUAGUUUUUAAAGAUGGAACUCCACCCUUUGCUUGGUUUUAAGUAUGUAUGGAAUGUUAUGAUAGGACAUAGUAGUAGCGGUGGUCAGACAUGGAAAUGGUGGGGAGACAAAAAUAUACAUGUGAAAUAAAACUCAGUAUUUUAAUAAAGUAGCACGGUUUCUAUUGA470- example_title: superoxide dismutase 1471 mask_index: 12472 mask_index_1based: 13473 masked_char: A474 output:475 - label: U476 score: 0.062892477 - label: Y478 score: 0.057641479 - label: W480 score: 0.05284481 - label: H482 score: 0.052837483 - label: C484 score: 0.052829485 pipeline_tag: fill-mask486 sequence_type: 3' UTR487 task: fill-mask488 text: ACAUUCCCUUGG<mask>UGUAGUCUGAGGCCCCUUAACUCAUCUGUUAUCCUGCUAGCUGUAGAAAUGUAUCCUGAUAAACAUUAAACACUGUAAUCUUAAAAGUGUAAUUGUGUGACUUUUUCAGAGUUGCUUUAAAGUACCUGUAGUGAGAAACUGAUUUAUGAUCACUUGGAAGAUUUGUAUAGUUUUAUAAAACUCAGUUAAAAUGUCUGUUUCAAUGACCUGUAUUUUGCCAGACUUAAAUCACAGAUGGGUAUUAAACUUGUCAGAAUUUCUUUGUCAUUCAAGCCUGUGAAUAAAAACCCUGUAUGGCACUUAUUAUGAGGCUAUUAAAAGAAUCCAAAUUCAAACUAAA489- example_title: hemoglobin subunit alpha 2490 mask_index: 13491 mask_index_1based: 14492 masked_char: A493 output:494 - label: G495 score: 0.32064496 - label: R497 score: 0.073522498 - label: S499 score: 0.068353500 - label: K501 score: 0.064896502 - label: V503 score: 0.042866504 pipeline_tag: fill-mask505 sequence_type: 3' UTR506 task: fill-mask507 text: CUGGAGCCUCGGU<mask>GCCGUUCCUCCUGCCCGCUGGGCCUCCCAACGGGCCCUCCUCCCCUCCUUGCACCGGCCCUUCCUGGUCUUUGAAUAAAGUCUGAGUGGGCAGCA508- example_title: BRAF proto-oncogene509 mask_index: 12510 mask_index_1based: 13511 masked_char: A512 output:513 - label: A514 score: 0.130968515 - label: R516 score: 0.091002517 - label: W518 score: 0.085338519 - label: D520 score: 0.077222521 - label: G522 score: 0.063232523 pipeline_tag: fill-mask524 sequence_type: 3' UTR525 task: fill-mask526 text: AACAAAUGAGUG<mask>GAGAGUUCAGGAGAGUAGCAACAAAAGGAAAAUAAAUGAACAUAUGUUUGCUUAUAUGUUAAAUUGAAUAAAAUACUCUCUUUUUUUUUAAGGUGAACCAAAGAACACUUGUGUGGUUAAAGACUAGAUAUAAUUUUUCCCCAAACUAAAAUUUAUACUUAACAUUGGAUUUUUAACAUCCAAGGGUUAAAAUACAUAGACAUUGCUAAAAAUUGGCAGAGCCUCUUCUAGAGGCUUUACUUUCUGUUCCGGGUUUGUAUCAUUCACUUGGUUAUUUUAAGUAGUAAACUUCAGUUUCUCAUGCAACUUUUGUUGCCAGCUAUCACAUGUCCACUAGGGACUCCAGAAGAAGACCCUACCUAUGCCUGUGUUUGCAGGUGAGAAGUUGGCAGUCGGUUAGCCUGGG527- example_title: H3 clustered histone 1528 mask_index: 17529 mask_index_1based: 18530 masked_char: A531 output:532 - label: A533 score: 0.06926534 - label: R535 score: 0.05768536 - label: W537 score: 0.052921538 - label: D539 score: 0.05124540 - label: M541 score: 0.048795542 pipeline_tag: fill-mask543 sequence_type: 3' UTR544 task: fill-mask545 text: UUACUGUGGUCUCUCUG<mask>CGGUCCAAGCAAAGGCUCUUUUCAGAGCCACCACCUUUUC546---547 548# SpliceBERT549 550Pre-trained model on messenger RNA precursor (pre-mRNA) using a masked language modeling (MLM) objective.551 552## Disclaimer553 554This is an UNOFFICIAL implementation of the [Self-supervised learning on millions of pre-mRNA sequences improves sequence-based RNA splicing prediction](https://doi.org/10.1101/2023.01.31.526427) by Ken Chen, et al.555 556The OFFICIAL repository of SpliceBERT is at [chenkenbio/SpliceBERT](https://github.com/chenkenbio/SpliceBERT).557 558> [!TIP]559> The MultiMolecule team has confirmed that the provided model and checkpoints are producing the same intermediate representations as the original implementation.560 561**The team releasing SpliceBERT did not write this model card for this model so this model card has been written by the MultiMolecule team.**562 563## Model Details564 565SpliceBERT is a [bert](https://huggingface.co/google-bert/bert-base-uncased)-style model pre-trained on a large corpus of messenger RNA precursor sequences in a self-supervised fashion. This means that the model was trained on the raw nucleotides of RNA sequences only, with an automatic process to generate inputs and labels from those texts. Please refer to the [Training Details](#training-details) section for more information on the training process.566 567### Variants568 569- **[multimolecule/splicebert](https://huggingface.co/multimolecule/splicebert)**: The SpliceBERT model.570- **[multimolecule/splicebert.510](https://huggingface.co/multimolecule/splicebert.510)**: The intermediate SpliceBERT model.571- **[multimolecule/splicebert-human.510](https://huggingface.co/multimolecule/splicebert-human.510)**: The intermediate SpliceBERT model pre-trained on human data only.572 573### Model Specification574 575<table>576<thead>577 <tr>578 <th>Variants</th>579 <th>Num Layers</th>580 <th>Hidden Size</th>581 <th>Num Heads</th>582 <th>Intermediate Size</th>583 <th>Num Parameters (M)</th>584 <th>FLOPs (G)</th>585 <th>MACs (G)</th>586 <th>Max Num Tokens</th>587 </tr>588</thead>589<tbody>590 <tr>591 <td>splicebert</td>592 <td rowspan="3">6</td>593 <td rowspan="3">512</td>594 <td rowspan="3">16</td>595 <td rowspan="3">2048</td>596 <td>19.72</td>597 <td>22.66</td>598 <td>11.27</td>599 <td>1024</td>600 </tr>601 <tr>602 <td><b>splicebert.510</b></td>603 <td rowspan="2">19.45</td>604 <td rowspan="2">22.56</td>605 <td rowspan="2">11.22</td>606 <td rowspan="2">510</td>607 </tr>608 <tr>609 <td>splicebert-human.510</td>610 </tr>611</tbody>612</table>613 614### Links615 616- **Code**: [multimolecule.splicebert](https://github.com/DLS5-Omics/multimolecule/tree/master/multimolecule/models/splicebert)617- **Data**: [UCSC Genome Browser](https://genome.ucsc.edu)618- **Paper**: [Self-supervised learning on millions of pre-mRNA sequences improves sequence-based RNA splicing prediction](https://doi.org/10.1101/2023.01.31.526427)619- **Developed by**: Ken Chen, Yue Zhou, Maolin Ding, Yu Wang, Zhixiang Ren, Yuedong Yang620- **Model type**: [BERT](https://huggingface.co/google-bert/bert-base-uncased)621- **Original Repository**: [chenkenbio/SpliceBERT](https://github.com/chenkenbio/SpliceBERT)622 623## Usage624 625The model file depends on the [`multimolecule`](https://multimolecule.danling.org) library. You can install it using pip:626 627```bash628pip install multimolecule629```630 631### Direct Use632 633#### Masked Language Modeling634 635You can use this model directly with a pipeline for masked language modeling:636 637```python638import multimolecule # you must import multimolecule to register models639from transformers import pipeline640 641predictor = pipeline("fill-mask", model="multimolecule/splicebert")642output = predictor("gguc<mask>cucugguuagaccagaucugagccu")643```644 645### Downstream Use646 647#### Extract Features648 649Here is how to use this model to get the features of a given sequence in PyTorch:650 651```python652from multimolecule import RnaTokenizer, SpliceBertModel653 654 655tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")656model = SpliceBertModel.from_pretrained("multimolecule/splicebert")657 658text = "UAGCUUAUCAGACUGAUGUUG"659input = tokenizer(text, return_tensors="pt")660 661output = model(**input)662```663 664#### Sequence Classification / Regression665 666> [!NOTE]667> This model is not fine-tuned for any specific task. You will need to fine-tune the model on a downstream task to use it for sequence classification or regression.668 669Here is how to use this model as backbone to fine-tune for a sequence-level task in PyTorch:670 671```python672import torch673from multimolecule import RnaTokenizer, SpliceBertForSequencePrediction674 675 676tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")677model = SpliceBertForSequencePrediction.from_pretrained("multimolecule/splicebert")678 679text = "UAGCUUAUCAGACUGAUGUUG"680input = tokenizer(text, return_tensors="pt")681label = torch.tensor([1])682 683output = model(**input, labels=label)684```685 686#### Token Classification / Regression687 688> [!NOTE]689> This model is not fine-tuned for any specific task. You will need to fine-tune the model on a downstream task to use it for token classification or regression.690 691Here is how to use this model as backbone to fine-tune for a nucleotide-level task in PyTorch:692 693```python694import torch695from multimolecule import RnaTokenizer, SpliceBertForTokenPrediction696 697 698tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")699model = SpliceBertForTokenPrediction.from_pretrained("multimolecule/splicebert")700 701text = "UAGCUUAUCAGACUGAUGUUG"702input = tokenizer(text, return_tensors="pt")703label = torch.randint(2, (len(text), ))704 705output = model(**input, labels=label)706```707 708#### Contact Classification / Regression709 710> [!NOTE]711> This model is not fine-tuned for any specific task. You will need to fine-tune the model on a downstream task to use it for contact classification or regression.712 713Here is how to use this model as backbone to fine-tune for a contact-level task in PyTorch:714 715```python716import torch717from multimolecule import RnaTokenizer, SpliceBertForContactPrediction718 719 720tokenizer = RnaTokenizer.from_pretrained("multimolecule/splicebert")721model = SpliceBertForContactPrediction.from_pretrained("multimolecule/splicebert")722 723text = "UAGCUUAUCAGACUGAUGUUG"724input = tokenizer(text, return_tensors="pt")725label = torch.randint(2, (len(text), len(text)))726 727output = model(**input, labels=label)728```729 730## Training Details731 732SpliceBERT used Masked Language Modeling (MLM) as the pre-training objective: taking a sequence, the model randomly masks 15% of the tokens in the input then runs the entire masked sentence through the model and has to predict the masked tokens. This is comparable to the Cloze task in language modeling.733 734### Training Data735 736The SpliceBERT model was pre-trained on messenger RNA precursor sequences from [UCSC Genome Browser](https://genome.ucsc.edu).737UCSC Genome Browser provides visualization, analysis, and download of comprehensive vertebrate genome data with aligned annotation tracks (known genes, predicted genes, ESTs, mRNAs, CpG islands, etc.).738 739SpliceBERT collected reference genomes and gene annotations from the UCSC Genome Browser for 72 vertebrate species. It applied [bedtools getfasta](https://bedtools.readthedocs.io/en/latest/content/tools/getfasta.html) to extract pre-mRNA sequences from the reference genomes based on the gene annotations. The pre-mRNA sequences are then used to pre-train SpliceBERT. The pre-training data contains 2 million pre-mRNA sequences with a total length of 65 billion nucleotides.740 741Note [`RnaTokenizer`][multimolecule.RnaTokenizer] will convert "T"s to "U"s for you, you may disable this behaviour by passing `replace_T_with_U=False`.742 743### Training Procedure744 745#### Preprocessing746 747SpliceBERT used masked language modeling (MLM) as the pre-training objective. The masking procedure is similar to the one used in BERT:748 749- Mask rate: 15%750- Replacement: `<mask>` for 80% of masked tokens751- Replacement: random token for 10% of masked tokens752- Replacement: unchanged token for 10% of masked tokens753 754#### Pre-training755 756The model was trained on 8 NVIDIA V100 GPUs.757 758- Optimizer: AdamW759- Learning rate: 1e-4760- Learning rate scheduler: ReduceLROnPlateau(patience=3)761 762SpliceBERT trained model in a two-stage training process:763 7641. Pre-train with sequences of a fixed length of 510 nucleotides.7652. Pre-train with sequences of a variable length between 64 and 1024 nucleotides.766 767The intermediate model after the first stage is available as `multimolecule/splicebert.510`.768 769SpliceBERT also pre-trained a model on human data only to validate the contribution of multi-species pre-training. The intermediate model after the first stage is available as `multimolecule/splicebert-human.510`.770 771## Citation772 773```bibtex774@article {chen2023self,775 author = {Chen, Ken and Zhou, Yue and Ding, Maolin and Wang, Yu and Ren, Zhixiang and Yang, Yuedong},776 title = {Self-supervised learning on millions of pre-mRNA sequences improves sequence-based RNA splicing prediction},777 elocation-id = {2023.01.31.526427},778 year = {2023},779 doi = {10.1101/2023.01.31.526427},780 publisher = {Cold Spring Harbor Laboratory},781 abstract = {RNA splicing is an important post-transcriptional process of gene expression in eukaryotic cells. Predicting RNA splicing from primary sequences can facilitate the interpretation of genomic variants. In this study, we developed a novel self-supervised pre-trained language model, SpliceBERT, to improve sequence-based RNA splicing prediction. Pre-training on pre-mRNA sequences from vertebrates enables SpliceBERT to capture evolutionary conservation information and characterize the unique property of splice sites. SpliceBERT also improves zero-shot prediction of variant effects on splicing by considering sequence context information, and achieves superior performance for predicting branchpoint in the human genome and splice sites across species. Our study highlighted the importance of pre-training genomic language models on a diverse range of species and suggested that pre-trained language models were promising for deciphering the sequence logic of RNA splicing.Competing Interest StatementThe authors have declared no competing interest.},782 URL = {https://www.biorxiv.org/content/early/2023/05/09/2023.01.31.526427},783 eprint = {https://www.biorxiv.org/content/early/2023/05/09/2023.01.31.526427.full.pdf},784 journal = {bioRxiv}785}786```787 788> [!NOTE]789> The artifacts distributed in this repository are part of the MultiMolecule project.790> If MultiMolecule supports your research, please cite the MultiMolecule project as follows:791 792```bibtex793@software{chen_2024_12638419,794 author = {Chen, Zhiyuan and Zhu, Sophia Y.},795 title = {MultiMolecule},796 doi = {10.5281/zenodo.12638419},797 publisher = {Zenodo},798 url = {https://doi.org/10.5281/zenodo.12638419},799 year = 2024,800 month = may,801 day = 4802}803```804 805## Contact806 807Please use GitHub issues of [MultiMolecule](https://github.com/DLS5-Omics/multimolecule/issues) for any questions or comments on the model card.808 809Please contact the authors of the [SpliceBERT paper](https://doi.org/10.1101/2023.01.31.526427) for questions or comments on the paper/model.810 811## License812 813This model implementation is licensed under the [GNU Affero General Public License](license.md).814 815For additional terms and clarifications, please refer to our [License FAQ](license-faq.md).816 817```spdx818SPDX-License-Identifier: AGPL-3.0-or-later819```