CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01slone /nllb-200-10M-sample Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.tabulartranslation1M<n<10M14 likes304 downloads3y agoHugging Face02vansh62 /nllb-200-10M-sample Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/vansh62/nllb-200-10M-sample.tabulartranslation1M<n<10M0 likes113 downloads13d agoHugging Face03vpermilp /nllb-200-distilled-600M-rust NLLB-200 This is the model card of NLLB-200's distilled 600M variant. Here are the metrics for that particular checkpoint. Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. Paper or other resource for more information NLLB Team et al, No… See the full description on the dataset page: https://huggingface.co/datasets/vpermilp/nllb-200-distilled-600M-rust.translation100K<n<1M1 likes108 downloads4y agoHugging Face04vpermilp /nllb-200-1.3B-rust NLLB-200 This is the model card of NLLB-200's 1.3B variant. Here are the metrics for that particular checkpoint. Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper. Paper or other resource for more information NLLB Team et al, No Language Left… See the full description on the dataset page: https://huggingface.co/datasets/vpermilp/nllb-200-1.3B-rust.4 likes93 downloads4y agoHugging Face05jasonrichdarmawan /nllb-200-6M-sample-embeddingOriginal dataset SONAR's author message What happens to the original dataset? Filter by blaser_sim >= 3.5 Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline What are the use cases? Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly tabular1M<n<10M0 likes53 downloads1y agoHugging Face06UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face07UMCU /pmc_patientcases_dutch_nllb200 Dataset Card for Pmc Patientcases Dutch Marianmt The original dataset was created by: Zhengyun Zhao, Qiao Jin, Fangyuan Chen, Tuorui Peng & Sheng Yu The source language: English The original data source. Data description Translation of PMC patient records using the NLLB200, using the PubScience library. For the translation we concatenated chunked translations to account for the maximum context length of the NLLB200 model. Note: these are unfiltered translations from an… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/pmc_patientcases_dutch_nllb200.0 likes14 downloads2y agoHugging Face08Youseff1987 /multilingual_translation_sft_nllb200_tokenizedtabular10K<n<100K0 likes14 downloads2y agoHugging Face09madiqbal /English_Hangul_NLLB200-3.3B0 likes13 downloads3y agoHugging Face10Slim205 /wiki_data_full_filtered_translated_nllb20000text10K<n<100K0 likes6 downloads2y agoHugging Face11Slim205 /hellaswag_ift_v02_filtered_translated_nllb20000text1K<n<10K0 likes5 downloads2y agoHugging Face12Slim205 /openbook_ift_v02_filtered_translated_nllb2000tabular1K<n<10K0 likes5 downloads2y agoHugging Face13Slim205 /hellaswag_ift_translated_nllb20000text10K<n<100K0 likes4 downloads2y agoHugging Face14Slim205 /sciq_ift_v02_filtered_translated_nllb2000text1K<n<10K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.