datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nllb-200-10M-sample
Dataset Card for "nllb-200-10M-sample"
This is a sample of nearly 10M sentence pairs from the NLLB-200
mined dataset allenai/nllb,
scored with the model facebook/blaser-2.0-qe
described in the SeamlessM4T paper.
The sample is not random; instead, we just took the top n sentence pairs from each translation direction.
The number n was computed with the goal of upsamping the directions that contain underrepresented languages.
Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/slone/nllb-200-10M-sample.nllb-200-10M-sample
Dataset Card for "nllb-200-10M-sample"
This is a sample of nearly 10M sentence pairs from the NLLB-200
mined dataset allenai/nllb,
scored with the model facebook/blaser-2.0-qe
described in the SeamlessM4T paper.
The sample is not random; instead, we just took the top n sentence pairs from each translation direction.
The number n was computed with the goal of upsamping the directions that contain underrepresented languages.
Nevertheless, the 187 languoids (language and script… See the full description on the dataset page: https://huggingface.co/datasets/vansh62/nllb-200-10M-sample.nllb-200-distilled-600M-rust
NLLB-200
This is the model card of NLLB-200's distilled 600M variant.
Here are the metrics for that particular checkpoint.
Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper.
Paper or other resource for more information NLLB Team et al, No… See the full description on the dataset page: https://huggingface.co/datasets/vpermilp/nllb-200-distilled-600M-rust.nllb-200-1.3B-rust
NLLB-200
This is the model card of NLLB-200's 1.3B variant.
Here are the metrics for that particular checkpoint.
Information about training algorithms, parameters, fairness constraints or other applied approaches, and features. The exact training algorithm, data and the strategies to handle data imbalances for high and low resource languages that were used to train NLLB-200 is described in the paper.
Paper or other resource for more information NLLB Team et al, No Language Left… See the full description on the dataset page: https://huggingface.co/datasets/vpermilp/nllb-200-1.3B-rust.nllb-200-6M-sample-embeddingOriginal dataset
SONAR's author message
What happens to the original dataset?
Filter by blaser_sim >= 3.5
Add new columns embedding1 and embedding2. The embeddings are generated by SONAR's TextToEmbeddingModelPipeline
What are the use cases?
Training a model with embeddings as input. For example, training a Sparse Autoencoder. This saves computation because we do not need to load the encoder during training. Also, we do not need to cache the encoder's output on-the-fly
apollo_english_guidelines_translated_to_dutch_with_nllb200
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the NLLB200-600M NTM.
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
pmc_patientcases_dutch_nllb200
Dataset Card for Pmc Patientcases Dutch Marianmt
The original dataset was created by: Zhengyun Zhao, Qiao Jin, Fangyuan Chen, Tuorui Peng & Sheng Yu
The source language: English
The original data source.
Data description
Translation of PMC patient records using the NLLB200, using the PubScience library.
For the translation we concatenated chunked translations to account for the maximum context length of the NLLB200 model.
Note: these are unfiltered translations from an… See the full description on the dataset page: https://huggingface.co/datasets/UMCU/pmc_patientcases_dutch_nllb200.multilingual_translation_sft_nllb200_tokenizedEnglish_Hangul_NLLB200-3.3Bwiki_data_full_filtered_translated_nllb20000hellaswag_ift_v02_filtered_translated_nllb20000openbook_ift_v02_filtered_translated_nllb2000hellaswag_ift_translated_nllb20000sciq_ift_v02_filtered_translated_nllb2000
