datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OriginalDatasetllmportuguese-dialects-ipa-synthetic
portuguese-dialects-ipa-synthetic
920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties
(European regional, insular, Brazilian regional, African/Asian/border national norms,
medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects),
Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho,
and Galician-Portuguese. Each row carries two IPA columns with distinct provenance.
Schema
sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.PORhatecheck-portuguese
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-portuguese.portuguese-legal-sentences-v0
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Contributions
@rufimelo99
If you use this work, please cite:
@InProceedings{MeloSemantic,
author="Melo, Rui
and Santos, Pedro A.
and Dias, Jo{\~a}o",
editor="Moniz, Nuno
and Vale, Zita
and… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/portuguese-legal-sentences-v0.PortBench-Market
PortBench Market Base Dataset
Dataset Description
A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation.
Asset Coverage
Asset Class
Instruments
Data Fields
Sources
Equities
126
OHLCV + return
Yahoo Finance (ETFs: broad market, sector, factor, international)
Bonds
16
Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.pexels-portrait
Pexels Portrait
This dataset was collected using the official Pexels API: https://www.pexels.com/api/
All fields returned by the API are recorded, except the source image URL. Only the original image URL is recorded, since other variants can be retrieved on-the-fly by modifying the URL parameters. The "alt" field can be used for simple text-to-image training/finetuning.
Image files are not provided. You need to download the image files by yourself. Example of downloading image… See the full description on the dataset page: https://huggingface.co/datasets/gaunernst/pexels-portrait.fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.r20-portfolio-ai-perception
Portfolio Interference in LLM Brand Perception (R20 to R21)
Supersession note: This dataset originally backed R20 (2026ab, superseded). R21 (2026ac, DOI 10.5281/zenodo.19765401) supersedes both R8 (2026q) and R20. R21 merges R8 theory with R20 empirical (9,925 obs across 40 brands, 13 models, 7 traditions) into a single analytical-empirical paper. New citations should reference Zharnikov (2026ac).
Dataset DOI: 10.57967/hf/8380
Current Paper (R21): 10.5281/zenodo.19765401 --… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r20-portfolio-ai-perception.Portfolio-Optimizationobligaciones-laborales-por-tamano
Obligaciones laborales por tamaño de empresa (España, 2026)
Tabla estructurada de las obligaciones laborales de las empresas españolas según su tamaño: las que aplican a todas las empresas desde el primer empleado (registro horario, horas extraordinarias, prevención de riesgos, protección de datos, protocolo frente al acoso, desconexión digital, calendario y vacaciones, teletrabajo si aplica), las que se activan a partir de 50 personas trabajadoras (canal de denuncias, plan de… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/obligaciones-laborales-por-tamano.PortugueseLegalSentences-v0
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Contributions
@rufimelo99
PortugueseLegalSentences-v3
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Extended version of rufimelo/PortugueseLegalSentences-v1
400000/50000/50000
Contributions
@rufimelo99
openHermes_portuguesediskos_qa
Dataset Card for DISKOS-QA
Dataset Summary
DISKOS-QA is an open benchmark for question answering in subsurface and petroleum-domain workflows. It was developed in the FORCE ecosystem and is built from public DISKOS-related oil and gas documents. The broader project uses a Neo4j knowledge graph, topic-based retrieval, Azure OpenAI models, and DeepEval-based filtering to generate and score high-quality question-answer pairs.
The public benchmark is distributed as a tabular… See the full description on the dataset page: https://huggingface.co/datasets/porestar/diskos_qa.aetherx-port-congestion-metrics
Aether-X Global Port Congestion Snapshot
Point-in-time snapshot of the Aether-X Port Congestion Oracle — predictive
congestion, ETA delay and freight-volatility signals for 15 of the world's
largest ports.
This static CSV is a frozen snapshot for research, backtesting and
dashboards. The live, continuously-updated signal is available through the
REST API and the Python SDK.
Files
port_metrics.csv — one row per port.
Schema
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/Aether-x/aetherx-port-congestion-metrics.synthetic_text_to_sql_th
Synthetic Text-to-SQL Thai Dataset
Thai translation of the gretelai/synthetic_text_to_sql dataset.
Dataset Description
This dataset contains Thai translations of synthetic text-to-SQL examples covering various domains and SQL patterns.
Source
Original Dataset: gretelai/synthetic_text_to_sql
Created by: Gretel.ai
Statistics
Split
Rows
Train
100,000
Test
5,851
Total
105,851
Columns
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/synthetic_text_to_sql_th.Portfolio-RebalancePatient-Message-Response-DraftingPaper: How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting (arxiv link)
Dataset Details:
The patient message response drafting dataset is designed to evaluate how well LLMs respond to patient messages in patient portal communication.
Each semi-synthetic patient message is paired with a real de-identified EHR from a patient at our collaborating hospital.
Each doctor response is written by a clinician, guided by clinician response themes… See the full description on the dataset page: https://huggingface.co/datasets/PortalPal-AI/Patient-Message-Response-Drafting.PortugueseLegalSentences-v1
Portuguese Legal Sentences
Collection of Legal Sentences from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for MLM and TSDAE
Contributions
@rufimelo99
portuguese_phonetic_lexicon
📚 Portuguese Phonetic Lexicon Dataset
This dataset contains phonetic and morphological information for Portuguese words, collected from the Portal da Língua Portuguesa. It was generated by scraping the site across multiple Portuguese-speaking regions and dialects.
🌍 Regional Coverage
The dataset includes words as spoken in ten regional variants:
🇵🇹 Lisbon (Standard and Non-Standard)
🇦🇴 Luanda
🇧🇷 Rio de Janeiro (Standard and Non-Standard)
🇧🇷 São Paulo (Standard… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese_phonetic_lexicon.Portuguese-Speech-Dataset
🎧 Portuguese Speech Dataset
The Portuguese Speech Dataset is a large-scale speech audio dataset designed to provide structured and high-quality audio data for modern AI and machine learning systems. It contains 195 hours of recorded speech data distributed across 894 files, available in MP3 and WAV formats, with a total size of 437 MB. This carefully curated audio dataset delivers diverse and representative voice data, with a balanced speaker distribution of 52% female and 48% male… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Portuguese-Speech-Dataset.maritime-vessel-schedule-port-congestion-coherence-risk-v0.1What this repo is for
Detect when vessel schedules stop matching port reality.
You use it to flag:
hidden delay risk when ETA looks stable but port signals collapse
upstream disruption when ETA slips without congestion signals
true congestion when yard and berth signals align with delay
Why it matters
Ports declare problems late.
Carriers update ETAs late.
Your users need the mismatch early.
Say “Next” and I’ll build the next port ops gap.
brazilian_portuguesespider_th
Spider Thai Dataset
Thai translation of the official Spider benchmark (A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task).
Dataset Description
This dataset contains Thai translations of the Spider text-to-SQL benchmark, translated from the official Spider data source.
Source
Original Dataset: Spider Benchmark
Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/spider_th.fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testrun001
fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testrun001
A curated registry of points of interest in downtown Portland, Oregon.
License
This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY).
You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source.
Contents
data.csv - sample data.
porto-seguro
porto-seguro
Created from AIOD platform
portuguese-hate-speech-superset
Portuguese Hate Speech Superset
This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.portable-power-station-specifications
Portable Power Station Specifications — 52 Models, Capacity/Output/Weight/Price (2026)
Manufacturer-published specifications for 52 portable power stations (Anker, Bluetti, EcoFlow, Goal Zero, Jackery) currently sold for home-backup and off-grid use: usable battery capacity (Wh), rated continuous and surge AC output (W), cell chemistry, weight, and approximate retail price. Includes derived Wh/lb (portability) and Wh/$ (value) for cross-brand comparison — 52 of 52 rows carry… See the full description on the dataset page: https://huggingface.co/datasets/rrhagentbiz/portable-power-station-specifications.africa-synth-trade-port-throughput-performance-all
African Port Throughput Performance | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: infrastructure_transport - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-trade-port-throughput-performance-all.
