CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes26k downloads4y agoHugging Face02openlanguagedata /flores_plusgated Dataset Card for FLORES+ FLORES+ is an evaluation benchmark dataset for multilingual machine translation. Dataset Details Dataset Description FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.tabulartext-generation100K<n<1M170 likes13k downloads2mo agoHugging Face03facebook /floresgated Dataset Card for Flores 200 Dataset Summary ⚠️ This repository is no longer being updated ⚠️ A newer version of the FLORES dataset managed by the Open Language Data Initiative is available at https://huggingface.co/datasets/openlanguagedata/flores_plus. FLORES is a benchmark dataset for machine translation between English and low-resource languages. The creation of FLORES-200 doubles the existing language coverage of FLORES-101. Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.tabulartext-generation1M<n<10M121 likes5.8k downloads4mo agoHugging Face04espnet /floras FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models. Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers. To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.audioautomatic-speech-recognition10K<n<100K15 likes3.8k downloads2mo agoHugging Face05mteb /flores FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English and low-resource languages. Task category t2t Domains Non-fiction, Encyclopaedic, Written Reference https://huggingface.co/datasets/facebook/flores How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["FloresBitextMining"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/flores.texttranslation1K<n<10K0 likes2.8k downloads1y agoHugging Face06severo /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M2 likes2.4k downloads4y agoHugging Face07facebook /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.tabulartranslation1K<n<10K2 likes1.8k downloads2y agoHugging Face08florencejiang /earnings25 Earnings25 A 500-hour speech benchmark for finance — S&P 500 earnings calls with reference transcripts, industry labels, and named-speaker attribution. Citation Earnings25 is introduced in our Interspeech 2026 paper, which sets out the sampling design, the evaluation protocol, and reference baselines for Whisper and Parakeet-TDT. Start there for the full picture. Jiang, D., Zhou, H., Wadhawan, A., Fahy, B., Ramesh, V., Weisberg, D., Derkachevskiy, D., Sheehan, H., Prasad, S., &… See the full description on the dataset page: https://huggingface.co/datasets/florencejiang/earnings25.audioautomatic-speech-recognitionn<1K2 likes1.1k downloads1mo agoHugging Face09haoranxu /FLORES-200text10K<n<100K3 likes727 downloads2y agoHugging Face10Kira-Floris /Afrivoice_Swahili_ASRaudio100K<n<1M0 likes695 downloads6mo agoHugging Face11florijanqosja /dibratext DibraText Dataset General Information Description: DibraText is an aggregated dataset consisting of Albanian language texts collected from various sources. It's designed for natural language processing tasks such as text classification, sentiment analysis, and machine learning model training. Author: Florijan Qosja Maintainer: Florijan Qosja (florijanqosja@gmail.com) Created: 27/04/2024 Last Updated: 27/04/2024 Language: Albanian (Language Code: sq) Volume: 100,378,107… See the full description on the dataset page: https://huggingface.co/datasets/florijanqosja/dibratext.text100M<n<1B0 likes455 downloads2y agoHugging Face12tartuNLP /flores-smugri-pairstabular10K<n<100K0 likes452 downloads1y agoHugging Face13yash9439 /flores200 FLORES-200 Subset (dev + devtest) This dataset contains the FLORES-200 multilingual machine translation dev and devtest splits in a consolidated Parquet format. It includes 997 dev examples and 1012 devtest examples across 200 languages, following the original FLORES-200 schema.Each row contains the same sentence translated into 200 language fields (e.g., eng_Latn, hin_Deva, zho_Hans, etc.). 📂 Dataset Structure Splits Split File # Examples… See the full description on the dataset page: https://huggingface.co/datasets/yash9439/flores200.texttranslation1K<n<10K0 likes411 downloads11mo agoHugging Face14VarunGumma /IGB_Flores_enxxtexttranslation10K<n<100K0 likes399 downloads1y agoHugging Face15APProjects /florida-layoffs-warn-act-notices-daily Florida WARN Act layoff notices — every filing we hold since 2015, one CSV, rebuilt daily 3,116 Florida WARN notices — every one this dataset holds, back to 2015 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16 · state source last checked 2026-09-24T14:05Z · official source: FloridaCommerce (REACT WARN list) — WARN notices. Florida employers must file a WARN Act notice with the state before a qualifying mass layoff or… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/florida-layoffs-warn-act-notices-daily.tabulartabular-classification1K<n<10K0 likes378 downloads2h agoHugging Face16Norod78 /OnceUponATime-florence2-captionsimagen<1K0 likes367 downloads2y agoHugging Face17VarunGumma /IGB_Flores_xxentexttranslation10K<n<100K0 likes328 downloads1y agoHugging Face18florin-hf /wiki_dump2018_no_duplicates Wikipedia Dump without Duplicates Dataset Summary This is a cleaned and de-duplicated version of the English Wikipedia dump dated December 20, 2018. Originally sourced from the DPR repository, it has been processed to remove duplicates, resulting in a final count of 20,970,784 passages, each consisting of 100 words. The original corpus is available for download via this link. The corpus is used in the research paper A Tale of Trust and Accuracy: Base vs. Instruct LLMs in… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/wiki_dump2018_no_duplicates.textquestion-answering10M<n<100M0 likes318 downloads2y agoHugging Face19mteb /FloresBitextMining FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English and low-resource languages. Task category t2t Domains Non-fiction, Encyclopaedic, Written Reference https://huggingface.co/datasets/facebook/flores Source datasets: mteb/flores How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/FloresBitextMining.texttranslation1K<n<10K1 likes304 downloads7mo agoHugging Face20aipicasso /soa-full-florence2 Smithsonian Open Access Dataset with Florence-2 Caption 日本語はこちら This dataset is made of soa-full. soa-full is an CC-0 image dataset from Smithsonian Open Access. However, the dataset does not contain the image caption. Therefore, we caption the images by Florence 2. Usage from datasets import load_dataset dataset = load_dataset("aipicasso/soa-full-florence2") Intended Use Research Vision & Language Develop text-to-image model or image-to-text model.… See the full description on the dataset page: https://huggingface.co/datasets/aipicasso/soa-full-florence2.imageimage-to-text1M<n<10M10 likes303 downloads2y agoHugging Face21deepearth /central-florida-native-plants DeepEarth Central Florida Native Plants Dataset v0.2.0 🌿 Dataset Summary A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research. Key Features 🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.tabularimage-classification10K<n<100K0 likes303 downloads1y agoHugging Face22hgissbkh /florestext100K<n<1M0 likes233 downloads5mo agoHugging Face23bri25yu /flores200_devtest_translation_pairs Dataset Card for "flores200_devtest_translation_pairs" More Information needed text10M<n<100M0 likes225 downloads3y agoHugging Face24Maximiliano-Flores-Dev /grok-demon-dataset-EStext1K<n<10K1 likes214 downloads8d agoHugging Face25DGME /FLORES-200texttranslation100K<n<1M0 likes208 downloads9mo agoHugging Face26meetsohail /translateplus-flores-benchmark TranslatePlus Translation Benchmark (FLORES 2026) This dataset contains benchmark results for TranslatePlus Translation API across 20 global languages using the FLORES dataset. 👉 Try the API: https://translateplus.io Methodology Dataset: FLORES (Facebook) Samples per language: 997 Source language: English Target languages: Top 20 global languages Evaluation metrics: BLEU (sacreBLEU) COMET (Unbabel/wmt22-comet-da) Evaluation type: Reference-based (human… See the full description on the dataset page: https://huggingface.co/datasets/meetsohail/translateplus-flores-benchmark.tabulartranslation10K<n<100K0 likes193 downloads6mo agoHugging Face27sfaamye /flores-all-configstabular100K<n<1M0 likes191 downloads3mo agoHugging Face28mbzuai-ugrip-statement-tuning /flores_101_instructiontext100K<n<1M0 likes183 downloads2y agoHugging Face29tomasmajercik /flores-parquet FLORES Parquet Parquet version of the FLORES dataset for efficient streaming. ⚠️ This is a derivative work: This dataset is a reformatted version of the original FLORES dataset created by Meta AI. All credit goes to the original authors. This version simply converts the data to Parquet format for easier streaming and usage. Usage from datasets import load_dataset # Load specific language ds = load_dataset("tomasmajercik/flores-parquet", name="fra_Latn"… See the full description on the dataset page: https://huggingface.co/datasets/tomasmajercik/flores-parquet.tabulartranslation100K<n<1M0 likes177 downloads9mo agoHugging Face30hlillemark /flores200_8_baseline Dataset Card for "flores200_8_baseline" More Information needed text10M<n<100M1 likes152 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.