datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
portuguese_benchmark
Portuguese Benchmark
This a collection of datasets in Portuguese initially meant to train and evaluate supervised language models such as BERT, RoBERTa, etc...
It contains 10 datasets and 18 Tasks for Classification (CLS), NLI, Semantic Similarity Scoring (STS) and Named-Entity Recognition (NER).
NER
Classification
NLI
STS
LeNER-Br
HateBR_offensive_binary
assin2-rte
assin2-sts
UlyssesNER-Br-PL-coarse
HateBR_offensive_level
UlyssesNER-Br-C-coarse… See the full description on the dataset page: https://huggingface.co/datasets/eduagarcia/portuguese_benchmark.education_data_portal_mirror_2026q3
Education Data Portal — Parquet Mirror (2026Q3 · Portal v0.26.1)
A complete mirror of the Urban Institute Education Data Portal datasets version 0.26.1, collected August 6, 2026, and converted from CSV to Apache Parquet format for efficient analytical use. Please note that the maintainers of this Huggingface Dataset have no affiliation with the Urban Institute or the Education Data Portal team.
Huge appreciation for all they do -- if you use this mirror, please make sure to… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror_2026q3.extraglue
This is the dataset card for extraGLUE.
You may be interested in some of the other datasets for Portuguese and in the models trained with them,
namely Albertina (encoders) and Gervásio (decoders) families.
ExtraGLUE
ExtraGLUE is a Portuguese dataset obtained by the automatic translation of some of the tasks in the GLUE and SuperGLUE benchmarks.
Two variants of Portuguese are considered, namely European Portuguese and American Portuguese.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/PORTULAN/extraglue.R1_Lite_take_and_place_the_portable_power_bank
R1_Lite_take_and_place_the_portable_power_bank
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: galaxea_r1_lite
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/R1_Lite_take_and_place_the_portable_power_bank.canonical-pores
Canonical Pores — Monte-Carlo Replay Packs
Converged Monte-Carlo diffusion walks in the three geometries that have exact analytical solutions —
parallel planes, cylinders, spheres — frozen so that any acquisition can be computed afterwards
without re-simulating.
600 substrates: 200 diameters per shape, 0.1–20.0 µm in 0.1 µm steps, each walked to
T = 200 ms at D₀ = 2.0×10⁻⁹ m²/s. One .rpk (safetensors) per substrate.
What you can replay
A replay pack is not a… See the full description on the dataset page: https://huggingface.co/datasets/SubstrateCommons/canonical-pores.SLR-Bench-Portuguese
🧠 SLR-Bench-Portuguese: Scalable Logical Reasoning Benchmark (Portuguese Edition)
SLR-Bench Multilingual Versions:
SLR-Bench-Portuguese is the Portuguese-language pendant of the original SLR-Bench dataset.
It follows the same symbolic structure, evaluation framework, and curriculum as the English version but provides all natural-language task prompts translated into Portuguese.
This enables systematic evaluation and training of Large Language Models… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/SLR-Bench-Portuguese.portal-frame-dataset
Formeo Portal Frame Dataset v2
50,000 verified samples / hour. Physics-verified structural samples from Formeo's Dataset Factory. This public sample is train-only input/output pairs and check labels under Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0). Generator, verifier, reasoning traces, and eval holdout stay proprietary. Need another structure type or engineering domain? Get in touch: info@formeo.ai.
Summary
Physics-verified… See the full description on the dataset page: https://huggingface.co/datasets/Formeo/portal-frame-dataset.wrist-usb-port-swap-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 896,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/choiwoong/wrist-usb-port-swap-test.chinese_porn_novelportuguese-eval-logs-olmo2-smollm3
Evaluation Logs on Portuguese Benchmarks for OLMo-2 and SmolLM3
These logs contain benchmark results across a suite of Portuguese-language tasks. The data consists of recordings of the performance of various 3 different models at different checkpoints throughout their pretraining runs:
SmolLM3
OLMo-2-0425-1B
OLMo-2-1124-7B
Splits
Each split (smollm3_3b, olmo2_1b, olmo2_7b) contains rows for model checkpoints and columns for benchmark scores (e.g., ASSIN2 RTE, ENEM… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/portuguese-eval-logs-olmo2-smollm3.PortBench-QA
PortBench QA Dataset
Dataset Description
6,269 structured question-answer pairs probing correlation-based financial reasoning for multi-asset portfolio management, generated from the PortBench Market Base Dataset.
Task Templates
Template
Task
Complexity
Pairs
T1
Return prediction — direction for next N days
1 (single asset)
1,000
T2
Risk assessment — VaR at given confidence level
1
1,000
T3
Position sizing — given max drawdown… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-QA.cot-portfoliowrist-usb-port-swap-test1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 896,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/choiwoong/wrist-usb-port-swap-test1.darwin-push-port-raw-07-09-2025_27-06-2026
Darwin Push Port (07 Sep 2025 – 27 Jun 2026)
Source
Apache Parquet files converted directly from raw XML archives hosted by ilovetrains.co.uk, which collects data via the Rail Data Marketplace Push Port.
Coverage & Caveats
Upstream Gaps: The dataset contains occasional hourly gaps inherited from the upstream archive.
Parsing Artifacts: While every effort has been made to preserve the raw XML faithfully, the automated conversion to JSON/Parquet may… See the full description on the dataset page: https://huggingface.co/datasets/TraderBeowulf/darwin-push-port-raw-07-09-2025_27-06-2026.education_data_portal_mirror
⚠️ Frozen Vintage — Education Data Portal Parquet Mirror (v0.24.0, February 2026)
This repository is frozen and will receive no further updates. It is preserved
permanently as a snapshot for reproducibility. For current data, use the
successor mirror:
brhkim/education_data_portal_mirror_2026q3
(Education Data Portal v0.26.1, collected August 2026).
If you are reproducing an analysis that originally used this mirror, you are in the
right place — keep your scripts pointed here.… See the full description on the dataset page: https://huggingface.co/datasets/brhkim/education_data_portal_mirror.ticuna-spanish-portuguese
Ticuna (tca) – Spanish – Portuguese Corpus
First text corpus for Ticuna (ISO 639-3 tca), a tonal language
isolate of the Brazil/Colombia/Peru tri-border.
Configs
| Config | Rows |
| parallel | train 43,248 / validation 596 / test 3,238 |
| monolingual | train 46,545 / validation 298 / test 1,613 |
| lexicon | train 10,419 / validation 568 / test 539 |
| instructions | train 52,836 |
| backtranslation | train 33,944 |
The short version of what matters… See the full description on the dataset page: https://huggingface.co/datasets/aimeri/ticuna-spanish-portuguese.hatecheck-portuguese
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.Emakhuwa-Portuguese-News-MT
News Parallel Dataset for Emakhuwa of Mozambique
This repository contains releases of parallel data for machine translation in Mozambican languages.
Currently, it supports one language pair, Portuguese-Emakhuwa, Emakhuwa being the widely spoken language in Mozambique.
Dataset Details
Dataset Description
Funded by: This dataset was created with support from Lacuna Fund, the world’s first collaborative effort to provide data scientists, researchers, and… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-Portuguese-News-MT.glue-ptptGLUE-PTPT is an European Portuguese translation of the GLUE benchmark using DeepL Pro.PortBench-Market
PortBench Market Base Dataset
Dataset Description
A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation.
Asset Coverage
Asset Class
Instruments
Data Fields
Sources
Equities
126
OHLCV + return
Yahoo Finance (ETFs: broad market, sector, factor, international)
Bonds
16
Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.ParlaSpeech-HR-benchmark_v1
ParlaSpeechHR Benchmark v1
A benchmark dataset of 17,083 Croatian parliamentary speech clips from ParlaSpeech-HR v1. Each clip includes audio (WAV) and rich metadata for speaker profiling tasks.
Contents
17,083 audio segments
Metadata: speaker info, party affiliation, birth year, gender
Word-level transcriptions and per-word start times (both raw and normalized variants)
Note: ~2,414 audio files referenced in the manifest are currently unavailable in this version… See the full description on the dataset page: https://huggingface.co/datasets/porupski/ParlaSpeech-HR-benchmark_v1.pexels-portrait
Pexels Portrait
This dataset was collected using the official Pexels API: https://www.pexels.com/api/
All fields returned by the API are recorded, except the source image URL. Only the original image URL is recorded, since other variants can be retrieved on-the-fly by modifying the URL parameters. The "alt" field can be used for simple text-to-image training/finetuning.
Image files are not provided. You need to download the image files by yourself. Example of downloading image… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/pexels-portrait.portal-rebot-pick-microcontroller
portal-rebot-pick-microcontroller
A LeRobot dataset of a 6-DOF arm picking a
microcontroller and placing it in a box. Episodes were collected by teleoperating a real
robot over the network with LiveKit Portal, using a
human-in-the-loop recording loop that aligns camera frames and joint state into clean,
synchronized trajectories.
Task: Pick the microcontroller and place it in the box
At a glance
Robot
Seeed reBot Arm B601-DM (seeed_b601_dm_follower)… See the full description on the dataset page: https://huggingface.co/datasets/binhpham/portal-rebot-pick-microcontroller.openbrush-portraits
OpenBrush Portraits
Every portrait painting from OpenBrush-75K — across all artists, movements, and centuries.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 13,059 you actually want.
Why this subset
Portraits across the full historical range — Renaissance bust portraits, Baroque chiaroscuro, Rococo society, Romantic, Realist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-portraits.ipfs_portugal_laws_ir
Portugal legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_portugal_laws (revision ``) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Portugal prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.
Primary key: entry_cid (CIDv1… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_portugal_laws_ir.Mr.Porter.Product.prices.Sweden
Mr Porter web scraped data
About the website
The EMEA region, specifically Sweden, has seen a significant rise in the luxury online retail industry, where Mr Porter operates. The growth has primarily been driven by the fast-paced digitalization, significant internet penetration, and a growing number of digitally native consumers. Additionally, Swedish consumers, renowned for their fashion-forward approach, have demonstrated a strong appetite for luxury fashion products… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Mr.Porter.Product.prices.Sweden.fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.PMR-Synth-Inboxesrobocasa365-PortionHotDogsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "PandaOmron",
"total_episodes": 504,
"total_frames": 424164,
"total_tasks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:504"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/robocasa365-PortionHotDogs.global-porphyry-copper-deposits
Quick Links
Website
GitHub
Hugging Face
LinkedIn
X
Global Porphyry Copper Deposits and Prospects, Modernized
2,394 porphyry copper deposits and prospects worldwide, compiled and
standardized from USGS's global database (Most up-to-date published information available to the authors as of Spring 2024).
Porphyry copper deposits are the world's primary source of copper and can
host secondary commodities USGS classifies as critical; their… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/global-porphyry-copper-deposits.
