scielo
Datasets
All datasets matching “scielo”scielo-clustering
SciELOClusteringP2P
Cluster Brazilian Portuguese scientific abstracts from the SciELO Brazil open-access library into 8 broad research areas (Health Sciences, Social Sciences, Agricultural Sciences, Biological/Life Sciences, Humanities & Arts, Engineering & Technology, Physical Sciences & Chemistry, Mathematics & Computer Science). Areas are consolidated from the Web-of-Science subject categories of each article; only pure CC-BY-4.0 articles are included.
Part of MTEB-BR — the… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/scielo-clustering.scielo_books
Dataset Summary
This dataset contains all text from open-access PDFs on scielo.org. As of Dec. 5 2021, the total number of books available is 962. Some of them are not in native PDF format (e.g. scanned images) though.
Supported Tasks and Leaderboards
sequence-modeling or language-modeling: The dataset can be used to train a language model.
Languages
As of Dec. 5 2021, there are 902 books in Portuguese, 55 in Spanish, and 5 in English.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/chenghao/scielo_books.scielo-summarization
LoRaLay: A Multilingual and Multimodal Dataset for Long Range and Layout-Aware Summarization
A collaboration between reciTAL, MLIA (ISIR, Sorbonne Université), Meta AI, and Università di Trento
SciELO dataset for summarization
SciELO is a dataset for summarization of research papers written in Spanish and Portuguese, for which layout information is provided.
Data Fields
article_id: article id
article_words: sequence of words constituting the body of… See the full description on the dataset page: https://huggingface.co/datasets/nglaura/scielo-summarization.scielo
Dataset Card for SciELO
Dataset Summary
A parallel corpus of full-text scientific articles collected from Scielo database in the following languages:English, Portuguese and Spanish.
The corpus is sentence aligned for all language pairs, as well as trilingual aligned for a small subset of sentences.
Alignment was carried out using the Hunalign algorithm.
Supported Tasks and Leaderboards
The underlying task is machine translation.
Languages
[More… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/scielo.scielo-corpus-csThis is a fork of ggg-llms-team/scielo-corpus.
This fork:
Adds checksums fields file_md5_checksum and file_sha256_checksum
Makes this dataset public under br-llm-data
SciELO-GL
Dataset Card for Spanish–Galician / English–Galician Scientific Corpus (SciELO)
Dataset Summary
The SciELO Corpus, hosted in theOPUS repository, is a large-scale parallel resource composed of full scientific articles extracted from the Scientific Electronic Library Online (SciELO). It provides high-quality sentence pairs between Spanish, Portuguese, and English across diverse academic domains.
This corpus is particularly valuable for training machine translation models… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/SciELO-GL.
