CoolFace
9 results

scielo

MTEB-BR /scielo-clustering SciELOClusteringP2P Cluster Brazilian Portuguese scientific abstracts from the SciELO Brazil open-access library into 8 broad research areas (Health Sciences, Social Sciences, Agricultural Sciences, Biological/Life Sciences, Humanities & Arts, Engineering & Technology, Physical Sciences & Chemistry, Mathematics & Computer Science). Areas are consolidated from the Web-of-Science subject categories of each article; only pure CC-BY-4.0 articles are included. Part of MTEB-BR — the… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/scielo-clustering.texttext-classification1K<n<10K0 likes1.1k downloads2mo agoHugging Facechenghao /scielo_books Dataset Summary This dataset contains all text from open-access PDFs on scielo.org. As of Dec. 5 2021, the total number of books available is 962. Some of them are not in native PDF format (e.g. scanned images) though. Supported Tasks and Leaderboards sequence-modeling or language-modeling: The dataset can be used to train a language model. Languages As of Dec. 5 2021, there are 902 books in Portuguese, 55 in Spanish, and 5 in English. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/scielo_books.textn<1K1 likes497 downloads4y agoHugging Facenglaura /scielo-summarization LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization A collaboration between reciTAL, MLIA (ISIR, Sorbonne Université), Meta AI, and Università di Trento SciELO dataset for summarization SciELO is a dataset for summarization of research papers written in Spanish and Portuguese, for which layout information is provided. Data Fields article_id: article id article_words: sequence of words constituting the body of… See the full description on the dataset page: https://huggingface.co/datasets/nglaura/scielo-summarization.summarization1 likes402 downloads3y agoHugging Facecommunity-datasets /scielo Dataset Card for SciELO Dataset Summary A parallel corpus of full-text scientific articles collected from Scielo database in the following languages:English, Portuguese and Spanish. The corpus is sentence aligned for all language pairs, as well as trilingual aligned for a small subset of sentences. Alignment was carried out using the Hunalign algorithm. Supported Tasks and Leaderboards The underlying task is machine translation. Languages [More… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/scielo.texttranslation1M<n<10M4 likes221 downloads2y agoHugging Facebr-llm-data /scielo-corpus-csgatedThis is a fork of ggg-llms-team/scielo-corpus. This fork: Adds checksums fields file_md5_checksum and file_sha256_checksum Makes this dataset public under br-llm-data text100K<n<1M0 likes95 downloads12d agoHugging Faceproxectonos /SciELO-GL Dataset Card for Spanish–Galician / English–Galician Scientific Corpus (SciELO) Dataset Summary The SciELO Corpus, hosted in theOPUS repository, is a large-scale parallel resource composed of full scientific articles extracted from the Scientific Electronic Library Online (SciELO). It provides high-quality sentence pairs between Spanish, Portuguese, and English across diverse academic domains. This corpus is particularly valuable for training machine translation models… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/SciELO-GL.texttranslation100K<n<1M2 likes61 downloads6mo agoHugging Face