datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.flores_plus
Dataset Card for FLORES+
FLORES+ is an evaluation benchmark dataset for multilingual machine translation.
Dataset Details
Dataset Description
FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.flores
Dataset Card for Flores 200
Dataset Summary
⚠️ This repository is no longer being updated ⚠️
A newer version of the FLORES dataset managed by the Open Language Data Initiative
is available at https://huggingface.co/datasets/openlanguagedata/flores_plus.
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
The creation of FLORES-200 doubles the existing language coverage of FLORES-101.
Given the nature of the new… See the full description on the dataset page: https://huggingface.co/datasets/facebook/flores.floras
FLORAS
FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language.
The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models.
Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers.
To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.flores
FloresBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
Task category
t2t
Domains
Non-fiction, Encyclopaedic, Written
Reference
https://huggingface.co/datasets/facebook/flores
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["FloresBitextMining"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/flores.flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.2M-Flores-ASL
2M-Flores
As part of the 2M-Belebele project, we have produced video recodings of ASL signing for all the dev and devtest
sentences in the original flores200 dataset.
To obtain ASL sign recordings, we provide translators of ASL and native signers with the English text version of the sentences to be recorded.
The interpreters are then asked to translate these sentences into ASL, create glosses for all sentences, and record their interpretations into ASL one sentence at a time.
The… See the full description on the dataset page: https://huggingface.co/datasets/facebook/2M-Flores-ASL.earnings25
Earnings25
A 500-hour speech benchmark for finance — S&P 500 earnings calls with reference
transcripts, industry labels, and named-speaker attribution.
Citation
Earnings25 is introduced in our Interspeech 2026 paper,
which sets out the sampling design, the evaluation protocol, and reference
baselines for Whisper and Parakeet-TDT. Start there for the full picture.
Jiang, D., Zhou, H., Wadhawan, A., Fahy, B., Ramesh, V., Weisberg, D.,
Derkachevskiy, D., Sheehan, H., Prasad, S., &… See the full description on the dataset page: https://huggingface.co/datasets/florencejiang/earnings25.FLORES-200Afrivoice_Swahili_ASRdibratext
DibraText Dataset
General Information
Description: DibraText is an aggregated dataset consisting of Albanian language texts collected from various sources. It's designed for natural language processing tasks such as text classification, sentiment analysis, and machine learning model training.
Author: Florijan Qosja
Maintainer: Florijan Qosja (florijanqosja@gmail.com)
Created: 27/04/2024
Last Updated: 27/04/2024
Language: Albanian (Language Code: sq)
Volume: 100,378,107… See the full description on the dataset page: https://huggingface.co/datasets/florijanqosja/dibratext.flores-smugri-pairsflores200
FLORES-200 Subset (dev + devtest)
This dataset contains the FLORES-200 multilingual machine translation dev and devtest splits in a consolidated Parquet format.
It includes 997 dev examples and 1012 devtest examples across 200 languages, following the original FLORES-200 schema.Each row contains the same sentence translated into 200 language fields (e.g., eng_Latn, hin_Deva, zho_Hans, etc.).
📂 Dataset Structure
Splits
Split
File
# Examples… See the full description on the dataset page: https://huggingface.co/datasets/yash9439/flores200.IGB_Flores_enxxflorida-layoffs-warn-act-notices-daily
Florida WARN Act layoff notices — every filing we hold since 2015, one CSV, rebuilt daily
3,116 Florida WARN notices — every one this dataset holds, back to 2015 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-16
· state source last checked 2026-09-24T14:05Z · official source: FloridaCommerce (REACT WARN list) — WARN notices.
Florida employers must file a WARN Act notice with the state before a qualifying
mass layoff or… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/florida-layoffs-warn-act-notices-daily.OnceUponATime-florence2-captionsIGB_Flores_xxenwiki_dump2018_no_duplicates
Wikipedia Dump without Duplicates
Dataset Summary
This is a cleaned and de-duplicated version of the English Wikipedia dump dated December 20, 2018. Originally sourced from the DPR repository, it has been processed to remove duplicates, resulting in a final count of 20,970,784 passages, each consisting of 100 words.
The original corpus is available for download via this link.
The corpus is used in the research paper A Tale of Trust and Accuracy: Base vs. Instruct LLMs in… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/wiki_dump2018_no_duplicates.FloresBitextMining
FloresBitextMining
An MTEB dataset
Massive Text Embedding Benchmark
FLORES is a benchmark dataset for machine translation between English and low-resource languages.
Task category
t2t
Domains
Non-fiction, Encyclopaedic, Written
Reference
https://huggingface.co/datasets/facebook/flores
Source datasets:
mteb/flores
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/FloresBitextMining.soa-full-florence2
Smithsonian Open Access Dataset with Florence-2 Caption
日本語はこちら
This dataset is made of soa-full.
soa-full is an CC-0 image dataset from Smithsonian Open Access. However, the dataset does not contain the image caption.
Therefore, we caption the images by Florence 2.
Usage
from datasets import load_dataset
dataset = load_dataset("aipicasso/soa-full-florence2")
Intended Use
Research Vision & Language
Develop text-to-image model or image-to-text model.… See the full description on the dataset page: https://huggingface.co/datasets/aipicasso/soa-full-florence2.central-florida-native-plants
DeepEarth Central Florida Native Plants Dataset v0.2.0
🌿 Dataset Summary
A comprehensive multimodal dataset featuring 33,665 observations of 232 native plant species from Central Florida. This dataset combines citizen science observations with state-of-the-art vision and language embeddings for advancing multimodal self-supervised ecological intelligence research.
Key Features
🌍 Spatiotemporal Coverage: Complete GPS coordinates and timestamps for all… See the full description on the dataset page: https://huggingface.co/datasets/deepearth/central-florida-native-plants.floresflores200_devtest_translation_pairs
Dataset Card for "flores200_devtest_translation_pairs"
More Information needed
grok-demon-dataset-ESFLORES-200translateplus-flores-benchmark
TranslatePlus Translation Benchmark (FLORES 2026)
This dataset contains benchmark results for TranslatePlus Translation API across 20 global languages using the FLORES dataset.
👉 Try the API: https://translateplus.io
Methodology
Dataset: FLORES (Facebook)
Samples per language: 997
Source language: English
Target languages: Top 20 global languages
Evaluation metrics:
BLEU (sacreBLEU)
COMET (Unbabel/wmt22-comet-da)
Evaluation type: Reference-based (human… See the full description on the dataset page: https://huggingface.co/datasets/meetsohail/translateplus-flores-benchmark.flores-all-configsflores_101_instructionflores-parquet
FLORES Parquet
Parquet version of the FLORES dataset for efficient streaming.
⚠️ This is a derivative work: This dataset is a reformatted version of the original FLORES dataset created by Meta AI. All credit goes to the original authors. This version simply converts the data to Parquet format for easier streaming and usage.
Usage
from datasets import load_dataset
# Load specific language
ds = load_dataset("tomasmajercik/flores-parquet", name="fra_Latn"… See the full description on the dataset page: https://huggingface.co/datasets/tomasmajercik/flores-parquet.flores200_8_baseline
Dataset Card for "flores200_8_baseline"
More Information needed
