CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01slone /nllb-200-10M-sample Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.tabulartranslation1M<n<10M14 likes335 downloads3y agoHugging Face02vansh62 /nllb-200-10M-sample Dataset Card for "nllb-200-10M-sample" This is a sample of nearly 10M sentence pairs from the NLLB-200 mined dataset allenai/nllb, scored with the model facebook/blaser-2.0-qe described in the SeamlessM4T paper. The sample is not random; instead, we just took the top n sentence pairs from each translation direction. The number n was computed with the goal of upsamping the directions that contain underrepresented languages. Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/vansh62/nllb-200-10M-sample.tabulartranslation1M<n<10M0 likes119 downloads13d agoHugging Face03jasonrichdarmawan /nllb-200-6M-sample-embeddingOriginal dataset SONAR's author message What happens to the original dataset? Filter by blaser_sim >= 3.5 Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline What are the use cases? Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly tabular1M<n<10M0 likes59 downloads1y agoHugging Face04UMCU /apollo_english_guidelines_translated_to_dutch_with_nllb200 Data description Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM. Acknowledgement The work received funding from the European Union's Horizon Europe research and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project). For more information on the background, see Datatools4Heart Huggingface/Website/Git tabulartext-generation10K<n<100K0 likes25 downloads2y agoHugging Face05Youseff1987 /multilingual_translation_sft_nllb200_tokenizedtabular10K<n<100K0 likes12 downloads2y agoHugging Face06Slim205 /wiki_data_full_filtered_translated_nllb20000text10K<n<100K0 likes6 downloads2y agoHugging Face07Slim205 /hellaswag_ift_v02_filtered_translated_nllb20000text1K<n<10K0 likes6 downloads2y agoHugging Face08Slim205 /openbook_ift_v02_filtered_translated_nllb2000tabular1K<n<10K0 likes6 downloads2y agoHugging Face09Slim205 /hellaswag_ift_translated_nllb20000text10K<n<100K0 likes4 downloads2y agoHugging Face10Slim205 /sciq_ift_v02_filtered_translated_nllb2000text1K<n<10K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.