datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nllb-200-10M-sample
Dataset Card for "nllb-200-10M-sample"
This is a sample of nearly 10M sentence pairs from the NLLB-200
mined dataset allenai/nllb,
scored with the model facebook/blaser-2.0-qe
described in the SeamlessM4T paper.
The sample is not random; instead, we just took the top n sentence pairs from each translation direction.
The number n was computed with the goal of upsamping the directions that contain underrepresented languages.
Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.nllb-200-10M-sample
Dataset Card for "nllb-200-10M-sample"
This is a sample of nearly 10M sentence pairs from the NLLB-200
mined dataset allenai/nllb,
scored with the model facebook/blaser-2.0-qe
described in the SeamlessM4T paper.
The sample is not random; instead, we just took the top n sentence pairs from each translation direction.
The number n was computed with the goal of upsamping the directions that contain underrepresented languages.
Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/vansh62/nllb-200-10M-sample.nllb-200-6M-sample-embeddingOriginal dataset
SONAR's author message
What happens to the original dataset?
Filter by blaser_sim >= 3.5
Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline
What are the use cases?
Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly
apollo_english_guidelines_translated_to_dutch_with_nllb200
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
multilingual_translation_sft_nllb200_tokenizedwiki_data_full_filtered_translated_nllb20000hellaswag_ift_v02_filtered_translated_nllb20000openbook_ift_v02_filtered_translated_nllb2000hellaswag_ift_translated_nllb20000sciq_ift_v02_filtered_translated_nllb2000
