CoolFace
Datasetpublic

latam-gpt/red_pajama_es_hq

RedPajama's High Quality Spanish subset What is this? The following is a high-quality dataset distilled from the Spanish subsection of RedPajama-Data-v2, created using the methodology proposed in FineWEB-Edu. Usage from datasets import load_dataset ds = load_dataset("latam-gpt/red_pajama_es_hq") Filtering by quality score Documents in this corpus are scored on academic quality from 2.5 to 5, with higher scores indicating better… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq.

sourceHugging Faceupdated 2y agoView on Hugging Face
11likes707downloads
18 commits on main
8073bef2y ago

Update README.md

ouhenio
f77f8162y ago

Update README.md

ouhenio
d328acd2y ago

Update README.md

ouhenio
813b31c2y ago

Update README.md

ouhenio
1cbcda22y ago

Update README.md

ouhenio
ebe4fb22y ago

Update README.md

ouhenio
384996b2y ago

Upload dataset (part 00002-of-00003)

tgomez
43825392y ago

Upload dataset (part 00001-of-00003)

tgomez
c7dcd602y ago

Upload dataset (part 00000-of-00003)

tgomez
88a52552y ago

Upload dataset (part 00002-of-00003)

tgomez
6cf1b022y ago

Upload dataset (part 00001-of-00003)

tgomez
2f947162y ago

Upload dataset (part 00000-of-00003)

tgomez
310f0042y ago

Upload dataset (part 00002-of-00003)

tgomez
efb2c882y ago

Upload dataset (part 00001-of-00003)

tgomez
3c23a6b2y ago

Upload dataset (part 00000-of-00003)

tgomez
9a9378d2y ago

Update README.md

ouhenio
98bf0352y ago

Create README.md

ouhenio
2c1b87d2y ago

initial commit

tgomez