costadev00/books-gutenberg-project-pt-br
Gutenberg Project TokenWeaver CPT 2048 - Unchunked This dataset contains full-document rows reconstructed from costadev00/gutenberg-project-tokenweaver-cpt-2048. The source dataset mixes reconstructed chunk sequences and singleton chunk rows. For this unchunked release, rows were grouped by metadata.id; when duplicated singleton rows were present for the same document, the reconstruction kept the series with the largest chunk_total. Text was joined with inferred text overlap.… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/books-gutenberg-project-pt-br.
07
No card is published for this repository, or it could not be fetched from Hugging Face right now.
