datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.flores_plus
Dataset Card for FLORES+
FLORES+ is an evaluation benchmark dataset for multilingual machine translation.
Dataset Details
Dataset Description
FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.flores
Dataset Card for Flores 200
Dataset Summary
⚠️ This repository is no longer being updated ⚠️
A newer version of the FLORES dataset managed by the Open Language Data Initiative
is available at https://huggingface.co/datasets/openlanguagedata/flores_plus.
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
The creation of FLORES-200 doubles the existing language coverage of FLORES-101.
Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.floras
FLORAS
FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language.
The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models.
Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers.
To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.flores
FloresBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
Task category
t2t
Domains
Non-fiction, Encyclopaedic, Written
Reference
https://huggingface.co/datasets/facebook/flores
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["FloresBitextMining"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/flores.flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.encoder-decoder-floresp-scoresflores200>The creation of FLORES200 doubles the existing language coverage of FLORES-101.
Given the nature of the new languages, which have less standardization and require
more specialized professional translations, the verification process became more complex.
This required modifications to the translation workflow. FLORES-200 has several languages
which were not translated from English. Specifically, several languages were translated
from Spanish, French, Russian and Modern Standard Arabic. Moreover, FLORES-200 also
includes two script alternatives for four languages. FLORES-200 consists of translations
from 842 distinct web articles, totaling 3001 sentences. These sentences are divided
into three splits: dev, devtest, and test (hidden). On average, sentences are approximately
21 words long.2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.IndicGenBench_flores_in
Dataset Card for Dataset Name
This repository contains the Flores-IN dataset released as a part of the paper "IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages"
Paper Link: https://arxiv.org/abs/2404.16816
Dataset Details
Overview
IndicGenBench is a multilingual, multi-way parallel benchmark for measuring language generation capabilities across diverse user-facing tasks in 29 Indic languages spanning 13… See the full description on the dataset page: https://huggingface.co/datasets/google/IndicGenBench_flores_in.earnings25
Earnings25
A 500-hour speech benchmark for finance — S&P 500 earnings calls with reference
transcripts, industry labels, and named-speaker attribution.
Citation
Earnings25 is introduced in our Interspeech 2026 paper,
which sets out the sampling design, the evaluation protocol, and reference
baselines for Whisper and Parakeet-TDT. Start there for the full picture.
Jiang, D., Zhou, H., Wadhawan, A., Fahy, B., Ramesh, V., Weisberg, D.,
Derkachevskiy, D., Sheehan, H., Prasad, S., &… See the full description on the dataset page: https://huggingface.co/datasets/florencejiang/earnings25.FLORES-200TroveLedger
🗃️ TroveLedger — Financial Time Series Dataset
A growing ledger of accumulated market history.
⚠️ Temporary Notice: Intraday Data Adjustments (January 2026)
What happened:A discrepancy has been identified in the minute- and hourly-resolution data: these series are currently not fully adjusted for stock splits and dividends. Daily-resolution data remains correctly adjusted (as provided by the source).
Why this matters:For accurate backtesting and model training –… See the full description on the dataset page: https://huggingface.co/datasets/Flori83/TroveLedger.Afrivoice_Swahili_ASRDCO_FLORES101_ALL_MODELS_EVAL_RESULTSflores200_baseline_all_mt5
Dataset Card for "flores200_baseline_all_mt5"
More Information needed
flores200_packing
Dataset Card for "flores200_packing"
More Information needed
dibratext
DibraText Dataset
General Information
Description: DibraText is an aggregated dataset consisting of Albanian language texts collected from various sources. It's designed for natural language processing tasks such as text classification, sentiment analysis, and machine learning model training.
Author: Florijan Qosja
Maintainer: Florijan Qosja (florijanqosja@gmail.com)
Created: 27/04/2024
Last Updated: 27/04/2024
Language: Albanian (Language Code: sq)
Volume: 100,378,107… See the full description on the dataset page: https://huggingface.co/datasets/florijanqosja/dibratext.flores-smugri-pairsflores200
FLORES-200 Subset (dev + devtest)
This dataset contains the FLORES-200 multilingual machine translation dev and devtest splits in a consolidated Parquet format.
It includes 997 dev examples and 1012 devtest examples across 200 languages, following the original FLORES-200 schema.Each row contains the same sentence translated into 200 language fields (e.g., eng_Latn, hin_Deva, zho_Hans, etc.).
📂 Dataset Structure
Splits
Split
File
# Examples… See the full description on the dataset page: https://huggingface.co/datasets/yash9439/flores200.IGB_Flores_enxxflores200_devtest_translation_pairs_mt5
Dataset Card for "flores200_devtest_translation_pairs_mt5"
More Information needed
flores200_packed2
Dataset Card for "flores200_packed2"
More Information needed
flores200_incomplete
Dataset Card for "flores200_incomplete"
More Information needed
florida-layoffs-warn-act-notices-daily
Florida WARN Act layoff notices — every filing we hold since 2015, one CSV, rebuilt daily
3,116 Florida WARN notices — every one this dataset holds, back to 2015 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16
· state source last checked 2026-09-24T14:05Z · official source: FloridaCommerce (REACT WARN list) — WARN notices.
Florida employers must file a WARN Act notice with the state before a qualifying
mass layoff or… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/florida-layoffs-warn-act-notices-daily.260217_OrthomosaikOnceUponATime-florence2-captionsflores200_packed2_mix_mt5
Dataset Card for "flores200_packed2_mix_mt5"
More Information needed
IGB_Flores_xxenwiki_dump2018_no_duplicates
Wikipedia Dump without Duplicates
Dataset Summary
This is a cleaned and de-duplicated version of the English Wikipedia dump dated December 20, 2018. Originally sourced from the DPR repository, it has been processed to remove duplicates, resulting in a final count of 20,970,784 passages, each consisting of 100 words.
The original corpus is available for download via this link.
The corpus is used in the research paper A Tale of Trust and Accuracy: Base vs. Instruct LLMs in… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/wiki_dump2018_no_duplicates.
