CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes26k downloads4y agoHugging Face02openlanguagedata /flores_plusgated Dataset Card for FLORES+ FLORES+ is an evaluation benchmark dataset for multilingual machine translation. Dataset Details Dataset Description FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.tabulartext-generation100K<n<1M170 likes13k downloads2mo agoHugging Face03facebook /floresgated Dataset Card for Flores 200 Dataset Summary ⚠️ This repository is no longer being updated ⚠️ A newer version of the FLORES dataset managed by the Open Language Data Initiative is available at https://huggingface.co/datasets/openlanguagedata/flores_plus. FLORES is a benchmark dataset for machine translation between English and low-resource languages. The creation of FLORES-200 doubles the existing language coverage of FLORES-101. Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.tabulartext-generation1M<n<10M121 likes5.8k downloads4mo agoHugging Face04espnet /floras FLORAS FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language. The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models. Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers. To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.audioautomatic-speech-recognition10K<n<100K15 likes3.8k downloads2mo agoHugging Face05mteb /flores FloresBitextMining An MTEB dataset Massive Text Embedding Benchmark FLORES is a benchmark dataset for machine translation between English and low-resource languages. Task category t2t Domains Non-fiction, Encyclopaedic, Written Reference https://huggingface.co/datasets/facebook/flores How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["FloresBitextMining"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/flores.texttranslation1K<n<10K0 likes2.8k downloads1y agoHugging Face06severo /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M2 likes2.4k downloads4y agoHugging Face07alvations /encoder-decoder-floresp-scores0 likes2.3k downloads2y agoHugging Face08Muennighoff /flores200>The creation of FLORES200 doubles the existing language coverage of FLORES-101. Given the nature of the new languages, which have less standardization and require more specialized professional translations, the verification process became more complex. This required modifications to the translation workflow. FLORES-200 has several languages which were not translated from English. Specifically, several languages were translated from Spanish, French, Russian and Modern Standard Arabic. Moreover, FLORES-200 also includes two script alternatives for four languages. FLORES-200 consists of translations from 842 distinct web articles, totaling 3001 sentences. These sentences are divided into three splits: dev, devtest, and test (hidden). On average, sentences are approximately 21 words long.translation23 likes2k downloads3y agoHugging Face09facebook /2M-Flores-ASL 2M-Flores As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest sentences in the original flores200 dataset. To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded. The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time. The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.tabulartranslation1K<n<10K2 likes1.8k downloads2y agoHugging Face10google /IndicGenBench_flores_in Dataset Card for Dataset Name This repository contains the Flores-IN dataset released as a part of the paper "IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages" Paper Link: https://arxiv.org/abs/2404.16816 Dataset Details Overview IndicGenBench is a multilingual, multi-way parallel benchmark for measuring language generation capabilities across diverse user-facing tasks in 29 Indic languages spanning 13… See the full description on the dataset page: https://huggingface.co/datasets/google/IndicGenBench_flores_in.translation10K<n<100K11 likes1.1k downloads2y agoHugging Face11florencejiang /earnings25 Earnings25 A 500-hour speech benchmark for finance — S&P 500 earnings calls with reference transcripts, industry labels, and named-speaker attribution. Citation Earnings25 is introduced in our Interspeech 2026 paper, which sets out the sampling design, the evaluation protocol, and reference baselines for Whisper and Parakeet-TDT. Start there for the full picture. Jiang, D., Zhou, H., Wadhawan, A., Fahy, B., Ramesh, V., Weisberg, D., Derkachevskiy, D., Sheehan, H., Prasad, S., &… See the full description on the dataset page: https://huggingface.co/datasets/florencejiang/earnings25.audioautomatic-speech-recognitionn<1K2 likes1.1k downloads1mo agoHugging Face12haoranxu /FLORES-200text10K<n<100K3 likes727 downloads2y agoHugging Face13Flori83 /TroveLedger 🗃️ TroveLedger — Financial Time Series Dataset A growing ledger of accumulated market history. ⚠️ Temporary Notice: Intraday Data Adjustments (January 2026) What happened:A discrepancy has been identified in the minute- and hourly-resolution data: these series are currently not fully adjusted for stock splits and dividends. Daily-resolution data remains correctly adjusted (as provided by the source). Why this matters:For accurate backtesting and model training –… See the full description on the dataset page: https://huggingface.co/datasets/Flori83/TroveLedger.time-series-forecastingn<1K0 likes711 downloads6mo agoHugging Face14Kira-Floris /Afrivoice_Swahili_ASRaudio100K<n<1M0 likes695 downloads6mo agoHugging Face15yileitu /DCO_FLORES101_ALL_MODELS_EVAL_RESULTS0 likes684 downloads3mo agoHugging Face16bri25yu /flores200_baseline_all_mt5 Dataset Card for "flores200_baseline_all_mt5" More Information needed 10M<n<100M0 likes545 downloads3y agoHugging Face17bri25yu /flores200_packing Dataset Card for "flores200_packing" More Information needed 1M<n<10M0 likes487 downloads4y agoHugging Face18florijanqosja /dibratext DibraText Dataset General Information Description: DibraText is an aggregated dataset consisting of Albanian language texts collected from various sources. It's designed for natural language processing tasks such as text classification, sentiment analysis, and machine learning model training. Author: Florijan Qosja Maintainer: Florijan Qosja (florijanqosja@gmail.com) Created: 27/04/2024 Last Updated: 27/04/2024 Language: Albanian (Language Code: sq) Volume: 100,378,107… See the full description on the dataset page: https://huggingface.co/datasets/florijanqosja/dibratext.text100M<n<1B0 likes455 downloads2y agoHugging Face19tartuNLP /flores-smugri-pairstabular10K<n<100K0 likes452 downloads1y agoHugging Face20yash9439 /flores200 FLORES-200 Subset (dev + devtest) This dataset contains the FLORES-200 multilingual machine translation dev and devtest splits in a consolidated Parquet format. It includes 997 dev examples and 1012 devtest examples across 200 languages, following the original FLORES-200 schema.Each row contains the same sentence translated into 200 language fields (e.g., eng_Latn, hin_Deva, zho_Hans, etc.). 📂 Dataset Structure Splits Split File # Examples… See the full description on the dataset page: https://huggingface.co/datasets/yash9439/flores200.texttranslation1K<n<10K0 likes411 downloads11mo agoHugging Face21VarunGumma /IGB_Flores_enxxtexttranslation10K<n<100K0 likes399 downloads1y agoHugging Face22hlillemark /flores200_devtest_translation_pairs_mt5 Dataset Card for "flores200_devtest_translation_pairs_mt5" More Information needed 10M<n<100M0 likes398 downloads3y agoHugging Face23bri25yu /flores200_packed2 Dataset Card for "flores200_packed2" More Information needed 10M<n<100M0 likes395 downloads4y agoHugging Face24bri25yu /flores200_incomplete Dataset Card for "flores200_incomplete" More Information needed 10M<n<100M0 likes394 downloads4y agoHugging Face25APProjects /florida-layoffs-warn-act-notices-daily Florida WARN Act layoff notices — every filing we hold since 2015, one CSV, rebuilt daily 3,116 Florida WARN notices — every one this dataset holds, back to 2015 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16 · state source last checked 2026-09-24T14:05Z · official source: FloridaCommerce (REACT WARN list) — WARN notices. Florida employers must file a WARN Act notice with the state before a qualifying mass layoff or… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/florida-layoffs-warn-act-notices-daily.tabulartabular-classification1K<n<10K0 likes378 downloads37m agoHugging Face26florianfischergeigergruppe /260217_Orthomosaikimagen<1K0 likes373 downloads7mo agoHugging Face27Norod78 /OnceUponATime-florence2-captionsimagen<1K0 likes367 downloads2y agoHugging Face28bri25yu /flores200_packed2_mix_mt5 Dataset Card for "flores200_packed2_mix_mt5" More Information needed 10M<n<100M0 likes332 downloads3y agoHugging Face29VarunGumma /IGB_Flores_xxentexttranslation10K<n<100K0 likes328 downloads1y agoHugging Face30florin-hf /wiki_dump2018_no_duplicates Wikipedia Dump without Duplicates Dataset Summary This is a cleaned and de-duplicated version of the English Wikipedia dump dated December 20, 2018. Originally sourced from the DPR repository, it has been processed to remove duplicates, resulting in a final count of 20,970,784 passages, each consisting of 100 words. The original corpus is available for download via this link. The corpus is used in the research paper A Tale of Trust and Accuracy: Base vs. Instruct LLMs in… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/wiki_dump2018_no_duplicates.textquestion-answering10M<n<100M0 likes318 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.