CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes26k downloads4y agoHugging Face02openlanguagedata /flores_plusgated Dataset Card for FLORES+ FLORES+ is an evaluation benchmark dataset for multilingual machine translation. Dataset Details Dataset Description FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.tabulartext-generation100K<n<1M170 likes13k downloads2mo agoHugging Face03facebook /floresgated Dataset Card for Flores 200 Dataset Summary ⚠️ This repository is no longer being updated ⚠️ A newer version of the FLORES dataset managed by the Open Language Data Initiative is available at https://huggingface.co/datasets/openlanguagedata/flores_plus. FLORES is a benchmark dataset for machine translation between English and low-resource languages. The creation of FLORES-200 doubles the existing language coverage of FLORES-101. Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.tabulartext-generation1M<n<10M121 likes5.8k downloads4mo agoHugging Face04severo /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M2 likes2.4k downloads4y agoHugging Face05tartuNLP /smugri-flores-testsetMultilingual FLORES-based benchmark for Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian. Expanded with Proper Karelian, Ludian, and Veps. Please, cite the following paper if you use Komi, Udmurt, Hill and Meadow Mari, Erzya, Moksha, Livonian, Mansi, and Livvi Karelian datasets: @inproceedings{ yankovskaya2023machine, title={Machine Translation for Low-resource Finno-Ugric Languages}, author={Lisa Yankovskaya and Maali Tars and Andre T{\"a}ttar and Mark… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/smugri-flores-testset.texttext-generation1K<n<10K4 likes125 downloads2y agoHugging Face06florath /coq-facts-props-proofs-gen0-v1 Dataset Name: Coq Facts, Propositions and Proofs Dataset Description The CoqFactsPropsProofs dataset aims to enhance Large Language Models' (LLMs) proficiency in interpreting and generating Coq code by providing a comprehensive collection of over 10,000 Coq source files. It encompasses a wide array of propositions, proofs, and definitions, enriched with metadata including source references and licensing information. This dataset is designed to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/florath/coq-facts-props-proofs-gen0-v1.texttext-generation100K<n<1M8 likes122 downloads3y agoHugging Face07Lots-of-LoRAs /task1514_flores_translation_entone Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1514_flores_translation_entone Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1514_flores_translation_entone.texttext-generation1K<n<10K0 likes118 downloads2y agoHugging Face08KBlueLeaf /danbooru2023-florence2-caption Danbooru2023 - Florence2 Caption dataset This dataset contains captions of danbooru2023 images generated by microsoft/Florence-2-large I use original one with task token Format parquet: key: the danbooru id of the image parsed: parsed florence 2 output of the image Stat MORE_DETAILED_CAPTION Entries: 7,438,449 Output Tokens (Min/Max/Mean/Median): Flan T5 Tokenizer: 19/736/120/114 DFN CLIP Tokenizer: 19/826/108.7/103 Qwen2 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-florence2-caption.texttext-to-image10M<n<100M66 likes112 downloads2y agoHugging Face09Maximiliano-Flores-Dev /agentlans-combined-roleplay_Dataset Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.texttext-generation1M<n<10M0 likes69 downloads4d agoHugging Face10Maximiliano-Flores-Dev /QuixiAI-dolphin_DatasetDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.texttext-generation1M<n<10M1 likes66 downloads4d agoHugging Face11Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes56 downloads4d agoHugging Face12Lots-of-LoRAs /task601_flores_translation_sntoen Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task601_flores_translation_sntoen Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task601_flores_translation_sntoen.texttext-generation1K<n<10K0 likes28 downloads2y agoHugging Face13lballore /llimba-flores-srd-eval LLiMba FLORES-200 Sardinian Evaluation Set A held-out evaluation set of 997 parallel sentences aligned across six languages (Sardinian, Italian, English, Spanish, French, Portuguese), derived from FLORES-200. Used to benchmark the LLiMba model's translation quality and reported in the LLiMba paper's BLEU and chrF tables. This is the exact evaluation set used to produce the published LLiMba translation results. Reproducing the paper's numbers requires this set plus… See the full description on the dataset page: https://huggingface.co/datasets/lballore/llimba-flores-srd-eval.texttranslationn<1K0 likes17 downloads5mo agoHugging Face14af5553g /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.text-generation0 likes16 downloads5mo agoHugging Face15Lots-of-LoRAs /task604_flores_translation_entosn Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task604_flores_translation_entosn Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task604_flores_translation_entosn.texttext-generation1K<n<10K0 likes15 downloads2y agoHugging Face16jeffbrian /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.text-generation0 likes14 downloads9mo agoHugging Face17FlorentMeyer /lpr_mnist LPR-MNIST We introduce LPR-MNIST, a synthetic dataset replicating the key aspects of syntax evolution on vehicle license plates, as explored in the paper Relaxed syntax modeling in Transformers for future-proof license plate recognition. LPR-MNIST is a collection of 100,000 synthetic image-text pairs, generated by concatenating 5 black-and-white digits from MNIST, each padded to 32x32. 32 pixels of zero-padding are then randomly shared between the left and right sides of the… See the full description on the dataset page: https://huggingface.co/datasets/FlorentMeyer/lpr_mnist.text-generation0 likes10 downloads1y agoHugging Face18GeoUpOrg /florian-hirt Florian Hirt Unternehmensprofile Dataset - Strukturierte Geschäftsdaten für KI-Systeme und Suchmaschinen. Branche: Dienstleistungen Standort: Wiesbaden, Deutschland Auf einen Blick Eigenschaft Wert Unternehmen Florian Hirt Branche Dienstleistungen Stadt Wiesbaden Land Deutschland Website https://florianhirt.de Telefon +49 1726 563984 E-Mail hirt@247365.de Über das Unternehmen Florian Hirt zählt zu den prägenden Strategen moderner… See the full description on the dataset page: https://huggingface.co/datasets/GeoUpOrg/florian-hirt.text-generationn<1K0 likes8 downloads7mo agoHugging Face19Yobitel /meta-llama-llama-3-1-70b-instruct__llm-mt-flores-200-mini-en-fr__019e3b92df93 meta-llama/Llama-3.1-70B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.8819 Chrf P50 0.9033 Chrf P95 1 TTFT P50 32.0341 ms Total P50 Ms 296.1723 Tokens Out Total 115 Run configuration Model: meta-llama/Llama-3.1-70B-Instruct @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-70b-instruct__llm-mt-flores-200-mini-en-fr__019e3b92df93.text-generationn<1K0 likes8 downloads4mo agoHugging Face20Yobitel /mistralai-mistral-7b-instruct-v0-3__llm-mt-flores-200-mini-en-fr__019e3b460885 mistralai/Mistral-7B-Instruct-v0.3 on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.767 Chrf P50 0.7445 Chrf P95 1 TTFT P50 11.9969 ms Total P50 Ms 154.8732 Tokens Out Total 181 Run configuration Model: mistralai/Mistral-7B-Instruct-v0.3 @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/mistralai-mistral-7b-instruct-v0-3__llm-mt-flores-200-mini-en-fr__019e3b460885.text-generationn<1K0 likes6 downloads4mo agoHugging Face21Yobitel /qwen-qwen2-5-7b-instruct__llm-mt-flores-200-mini-en-fr__019e3b39c5bb Qwen/Qwen2.5-7B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.8912 Chrf P50 0.9439 Chrf P95 1 TTFT P50 14.1119 ms Total P50 Ms 115.5118 Tokens Out Total 114 Run configuration Model: Qwen/Qwen2.5-7B-Instruct @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB HBM3 Driver:… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-7b-instruct__llm-mt-flores-200-mini-en-fr__019e3b39c5bb.text-generationn<1K0 likes5 downloads4mo agoHugging Face22Yobitel /google-gemma-2-9b-it__llm-mt-flores-200-mini-en-fr__019e3b9789b8 google/gemma-2-9b-it on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.7462 Chrf P50 0.6872 Chrf P95 1 TTFT P50 16.7669 ms Total P50 Ms 291.575 Tokens Out Total 199 Run configuration Model: google/gemma-2-9b-it @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB HBM3 Driver: 580.126.09… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-mt-flores-200-mini-en-fr__019e3b9789b8.text-generationn<1K0 likes4 downloads4mo agoHugging Face23Yobitel /meta-llama-llama-3-1-8b-instruct__llm-mt-flores-200-mini-en-fr__019e3b307ea3 meta-llama/Llama-3.1-8B-Instruct on llm.mt.flores-200-mini-en-fr (NVIDIA H100 80GB HBM3) Back to leaderboard Headline metrics Metric Value Unit N Samples 8 N Ok 8 Ok Rate 1 Chrf Mean 0.7749 Chrf P50 0.7614 Chrf P95 0.96 TTFT P50 14.1551 ms Total P50 Ms 140.4766 Tokens Out Total 151 Run configuration Model: meta-llama/Llama-3.1-8B-Instruct @ unknown00 Engine: vllm vunknown Quantization: fp16 Hardware: NVIDIA H100 80GB… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/meta-llama-llama-3-1-8b-instruct__llm-mt-flores-200-mini-en-fr__019e3b307ea3.text-generationn<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.