CoolFace
Datasetpublicgated

costadev00/books-gutenberg-project-pt-br

Gutenberg Project TokenWeaver CPT 2048 - Unchunked This dataset contains full-document rows reconstructed from costadev00/gutenberg-project-tokenweaver-cpt-2048. The source dataset mixes reconstructed chunk sequences and singleton chunk rows. For this unchunked release, rows were grouped by metadata.id; when duplicated singleton rows were present for the same document, the reconstruction kept the series with the largest chunk_total. Text was joined with inferred text overlap.… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/books-gutenberg-project-pt-br.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes5downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
costadev00/books-gutenberg-project-pt-br · CoolFace