CoolFace
20 results

latam

LatamBoard /leaderboard-results mauroibz/leaderboard-results Results from model evaluations on the leaderboard This dataset contains evaluation results from the leaderboard system. Structure Each JSON file contains results for a specific model evaluation Files are organized by organization/model structure Each result file includes: Model configuration Evaluation results across different benchmarks Metadata about the evaluation run Usage These results are used by the… See the full description on the dataset page: https://huggingface.co/datasets/LatamBoard/leaderboard-results.0 likes1k downloads3mo agoHugging Facelatam-gpt /red_pajama_es_hq RedPajama's High Quality Spanish subset What is this? The following is a high-quality dataset distilled from the Spanish subsection of RedPajama-Data-v2, created using the methodology proposed in FineWEB-Edu. Usage from datasets import load_dataset ds = load_dataset("latam-gpt/red_pajama_es_hq") Filtering by quality score Documents in this corpus are scored on academic quality from 2.5 to 5, with higher scores indicating better quality. The… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/red_pajama_es_hq.tabular100M<n<1B11 likes623 downloads2y agoHugging Facelatam-gpt /fineweb2-spa_Latn-edutabular100M<n<1B1 likes518 downloads2y agoHugging Facelatam-gpt /LatamGPT-Corpus-1.0gated LatamGPT-Corpus-1.0 🌐 Language versions: English | Español | Português 🔗 Project links: Official LatamGPT website | Corpus dashboard 🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0. Dataset description Summary LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.imagetext-generation100M<n<1B8 likes378 downloads10d agoHugging FaceGianDiego /latam-spanish-speech-orpheus-tts-24khz LATAM Spanish High-Quality Speech Dataset (24kHz - Orpheus TTS Ready) Dataset Description This dataset contains approximately 24 hours of high-quality speech audio in Latin American Spanish, specifically prepared for Text-to-Speech (TTS) applications like OrpheusTTS, which require a 24kHz sampling rate. The audio files are derived from the Crowdsourced high-quality speech datasets made by Google and were obtained via OpenSLR. The original recordings were high-quality… See the full description on the dataset page: https://huggingface.co/datasets/GianDiego/latam-spanish-speech-orpheus-tts-24khz.audiotext-to-speech10K<n<100K16 likes333 downloads1y agoHugging Facelatam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K8 likes270 downloads2mo agoHugging Face