CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes27k downloads4y agoHugging Face02openlanguagedata /flores_plusgated Dataset Card for FLORES+ FLORES+ is an evaluation benchmark dataset for multilingual machine translation. Dataset Details Dataset Description FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.tabulartext-generation100K<n<1M170 likes13k downloads2mo agoHugging Face03facebook /floresgated Dataset Card for Flores 200 Dataset Summary ⚠️ This repository is no longer being updated ⚠️ A newer version of the FLORES dataset managed by the Open Language Data Initiative is available at https://huggingface.co/datasets/openlanguagedata/flores_plus. FLORES is a benchmark dataset for machine translation between English and low-resource languages. The creation of FLORES-200 doubles the existing language coverage of FLORES-101. Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.tabulartext-generation1M<n<10M121 likes5.8k downloads4mo agoHugging Face04severo /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M2 likes2.4k downloads4y agoHugging Face05tartuNLP /smugri-flores-testsetMultilingual FLORES-based benchmark for Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian. Expanded with Proper Karelian, Ludian, and Veps. Please, cite the following paper if you use Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian datasets: @inproceedings{ yankovskaya2023machine, title={Machine Translation for Low-resource Finno-Ugric Languages}, author={Lisa Yankovskaya and Maali Tars and Andre T{\"a}ttar and Mark… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri-flores-testset.texttext-generation1K<n<10K4 likes126 downloads2y agoHugging Face06Lots-of-LoRAs /task1514_flores_translation_entone Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1514_flores_translation_entone Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1514_flores_translation_entone.texttext-generation1K<n<10K0 likes101 downloads2y agoHugging Face07Maximiliano-Flores-Dev /QuixiAI-dolphin_DatasetDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.texttext-generation1M<n<10M1 likes82 downloads4d agoHugging Face08Maximiliano-Flores-Dev /agentlans-combined-roleplay_Dataset Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.texttext-generation1M<n<10M0 likes70 downloads4d agoHugging Face09Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes60 downloads4d agoHugging Face10Lots-of-LoRAs /task601_flores_translation_sntoen Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task601_flores_translation_sntoen Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task601_flores_translation_sntoen.texttext-generation1K<n<10K0 likes28 downloads2y agoHugging Face11lballore /llimba-flores-srd-eval LLiMba FLORES-200 Sardinian Evaluation Set A held-out evaluation set of 997 parallel sentences aligned across six languages (Sardinian, Italian, English, Spanish, French, Portuguese), derived from FLORES-200. Used to benchmark the LLiMba model's translation quality and reported in the LLiMba paper's BLEU and chrF tables. This is the exact evaluation set used to produce the published LLiMba translation results. Reproducing the paper's numbers requires this set plus… See the full description on the dataset page: https://huggingface.co/datasets/lballore/llimba-flores-srd-eval.texttranslationn<1K0 likes17 downloads5mo agoHugging Face12af5553g /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.text-generation0 likes17 downloads5mo agoHugging Face13Lots-of-LoRAs /task604_flores_translation_entosn Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task604_flores_translation_entosn Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task604_flores_translation_entosn.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face14jeffbrian /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.text-generation0 likes15 downloads9mo agoHugging Face15Yobitel /meta-llama-llama-3-1-70b-instruct__llm-mt-flores-200-mini-en-fr__019e3b92df93 meta-llama/Llama-3.1-70B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.8819 Chrf P50 0.9033 Chrf P95 1 TTFT P50 32.0341 ms Total P50 Ms 296.1723 Tokens Out Total 115 Run configuration Model: meta-llama/Llama-3.1-70B-Instruct @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-70b-instruct__llm-mt-flores-200-mini-en-fr__019e3b92df93.text-generationn<1K0 likes8 downloads4mo agoHugging Face16Yobitel /mistralai-mistral-7b-instruct-v0-3__llm-mt-flores-200-mini-en-fr__019e3b460885 mistralai/Mistral-7B-Instruct-v0.3 on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.767 Chrf P50 0.7445 Chrf P95 1 TTFT P50 11.9969 ms Total P50 Ms 154.8732 Tokens Out Total 181 Run configuration Model: mistralai/Mistral-7B-Instruct-v0.3 @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/mistralai-mistral-7b-instruct-v0-3__llm-mt-flores-200-mini-en-fr__019e3b460885.text-generationn<1K0 likes6 downloads4mo agoHugging Face17Yobitel /qwen-qwen2-5-7b-instruct__llm-mt-flores-200-mini-en-fr__019e3b39c5bb Qwen/Qwen2.5-7B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.8912 Chrf P50 0.9439 Chrf P95 1 TTFT P50 14.1119 ms Total P50 Ms 115.5118 Tokens Out Total 114 Run configuration Model: Qwen/Qwen2.5-7B-Instruct @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB HBM3 Driver:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-7b-instruct__llm-mt-flores-200-mini-en-fr__019e3b39c5bb.text-generationn<1K0 likes5 downloads4mo agoHugging Face18Yobitel /google-gemma-2-9b-it__llm-mt-flores-200-mini-en-fr__019e3b9789b8 google/gemma-2-9b-it on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.7462 Chrf P50 0.6872 Chrf P95 1 TTFT P50 16.7669 ms Total P50 Ms 291.575 Tokens Out Total 199 Run configuration Model: google/gemma-2-9b-it @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB HBM3 Driver: 580.126.09… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-mt-flores-200-mini-en-fr__019e3b9789b8.text-generationn<1K0 likes4 downloads4mo agoHugging Face19Yobitel /meta-llama-llama-3-1-8b-instruct__llm-mt-flores-200-mini-en-fr__019e3b307ea3 meta-llama/Llama-3.1-8B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.7749 Chrf P50 0.7614 Chrf P95 0.96 TTFT P50 14.1551 ms Total P50 Ms 140.4766 Tokens Out Total 151 Run configuration Model: meta-llama/Llama-3.1-8B-Instruct @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-mt-flores-200-mini-en-fr__019e3b307ea3.text-generationn<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.