datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soma-competition-datasetSomali-ASR-Subset-68H
Somali ASR Subset 68H
Somali speech dataset for automatic speech recognition.
vr_m3_soma_retargeteashdatasets
Dataset Card for "ashdatasets"
More Information needed
Nordland
Nordland Dataset
This dataset is from the original videos released here: https://nrkbeta.no/2013/01/15/nordlandsbanen-minute-by-minute-season-by-season/
Citation Information
Please cite the original publication if you use this dataset.
Sünderhauf, Niko, Peer Neubert, and Peter Protzel. "Are we there yet? Challenging SeqSLAM on a 3000 km journey across all four seasons." Proc. of Workshop on Long-Term Autonomy, IEEE International Conference on Robotics and Automation… See the full description on the dataset page: https://huggingface.co/datasets/Somayeh-h/Nordland.DocParsingBenchDocParsingBench is a document intelligence benchmark of 1,400 images, systematically collected and annotated from real business workflows. It is the first dataset to systematically catalogue the document elements most frequently encountered in enterprise settings, covering five major domains: finance, legal, scientific research, manufacturing, and education.
🆕 Latest Updates
[2026.04.17] DocParsingBench evaluation toolkit release. Unified scoring is now available for the three… See the full description on the dataset page: https://huggingface.co/datasets/SoMarkAI/DocParsingBench.soma-to-multiple-robots
Thank you, Lambda
We thank Lambda for the compute support behind this multi-robot motion-generation and simulation effort.
SOMA to Multiple Robots
AlphaMotion is our self-developed cross-embodiment motion model. This repository collects robot-specific motion references generated from SOMA, visual comparisons, and saved downstream simulation evidence. It currently covers Unitree H2, Zhiyuan A3, and Tiangong. AlphaMotion is the model name; H2, A3 and… See the full description on the dataset page: https://huggingface.co/datasets/sajio/soma-to-multiple-robots.anv-data-ke-somali-fullanv-data-ke-somali-fullsoma-umr-release_v260714_a2m_52kyahoo-finance-data
The Financial data from Yahoo!
*** Key Points to Note ***
All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes.
I will update the data regularly, and you are welcome to follow this project and use the data.
Each time the data is updated, I will record the update time in spec.json.
Data Usage Instructions
Use DuckDB or… See the full description on the dataset page: https://huggingface.co/datasets/somamohanty/yahoo-finance-data.RoMo-SOMA-77
RoMo-SOMA-77 — RoMo Body+Hand Motion in 933-D Kimodo SOMA-77 Features
RoMo-SOMA-77 is the RoMo body+hand corpus packed in a 933-dimensional Kimodo SOMA-77 motion-feature representation, paired with rich multi-level text descriptions. It is the publication target for the SOMA-based body-and-hand model family.
Scope: paper-core (romo_official = True), matching RoMo-SMPL, RoMo-HML-263, and RoMo-272. A small number of clips are dropped where SOMA conversion produced non-finite… See the full description on the dataset page: https://huggingface.co/datasets/RoMoDataset/RoMo-SOMA-77.spending-archive
SomaliScan: US Government Spending Archive (2003–2026)
A unified, public-domain archive of US government spending, campaign
finance, lobbying, and federal employment data — aggregated from public
records into a single queryable corpus.
60 tables · ~696M rows · ~37 GB compressed Parquet · CC0 1.0
Quickstart
Every table is Apache Parquet. The fastest way to use this dataset is
DuckDB — install it once, then query directly
from this dataset without downloading… See the full description on the dataset page: https://huggingface.co/datasets/somaliscan/spending-archive.Somaliipfs_somalia_laws
Somalia Federal Laws and Constitution (parliament.gov.so / moj.gov.so)
Research snapshot of official national legislation from Federal Parliament (parliament.gov.so) + Ministry of Justice and Constitutional Affairs (moj.gov.so).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-18
Coverage
catalog-backed incomplete (parliament.gov.so WP media Sharci/Dastuur PDFs live;… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_somalia_laws.34minessomaliweb-v1
SomaliWeb v1 — Quality-filtered Somali web corpus
📄 Paper: arXiv:2605.18232 — SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
💻 Construction pipeline (MIT): github.com/khaledyusuf44/somali-corpus
SomaliWeb v1 is a cleaned, deduplicated, and quality-filtered Somali-language web corpus of ~303 million tokens (819,322 documents), built by aggregating three public Somali-heavy web distributions (HPLT v2… See the full description on the dataset page: https://huggingface.co/datasets/khaledyusuf44/somaliweb-v1.roll_into_tape_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/somasekharkakarla/roll_into_tape_v1.Mohamed-diirowsomali-multilingual-infopankki
Somali Multilingual Infopankki
somali-multilingual-infopankki is a parallel corpus containing multilingual translation pairs that involve the Somali (so) language. This dataset has been filtered and extracted from the original Helsinki-NLP/opus_infopankki corpus.
It is designed to support machine translation (NMT), multilingual sentence alignment, and Somali natural language processing (NLP) research.
Dataset Details
Source Dataset: Helsinki-NLP/opus_infopankki… See the full description on the dataset page: https://huggingface.co/datasets/tufaax/somali-multilingual-infopankki.ipfs_somalia_laws_ir
Somalia legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_somalia_laws (revision ac6aefed0dbe007940d49b85d19917a98d2f8c0d) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Somalia prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_somalia_laws_ir.somali-combined-asr-stt-dataset
Somali Combined ASR/STT Dataset
A Somali automatic speech recognition (ASR) / speech-to-text (STT) dataset combining
synthetic TTS-generated audio and other Somali speech sources, deduplicated by transcript
text and split into train/validation/test.
Dataset Summary
Language: Somali (so)
Task: Automatic Speech Recognition / Speech-to-Text
Audio format: WAV, 16 kHz mono
Total examples: 8,226 (after deduplication)
Total audio: ~6 hours
Split
Examples
Audio… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimDayax/somali-combined-asr-stt-dataset.Adam-pro-soma-retargetyahoo-finance-data
The Financial data from Yahoo!
*** Key Points to Note ***
All financial data is sourced from Yahoo!Ⓡ Finance, Nasdaq!Ⓡ, and the U.S. Department of the Treasury via publicly available APIs, and is intended for research and educational purposes.
I will update the data regularly, and you are welcome to follow this project and use the data.
Each time the data is updated, I will record the update time in spec.json.
Data Usage Instructions
Use DuckDB or… See the full description on the dataset page: https://huggingface.co/datasets/somAzzz/yahoo-finance-data.somali-100k-saaxiib-conversations
🇸🇴 Somali 100K AI Friend Conversations (Saaxiib AI)
The largest, cleanest, and most emotionally aware Somali conversational dataset ever built. Designed specifically to align language models into authentic, empathetic, and witty Somali AI Companions & Friends rather than dry informative tutors.
🌟 Key Characteristics
Intent-Locked Empathy: Zero emotional mismatch. Fatigue receives rest comfort, debt disputes receive financial advice, celebrations receive shared… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-saaxiib-conversations.somali-100k-saaxiib-conversations-v2
🇸🇴 Somali 100K Multi-Turn Saaxiib AI Dataset v2
This dataset contains 100,000 Extended Multi-Turn Dialogues (530,188 total turns) designed to train conversational AI companions in native spoken Somali.
🌟 Key Improvements in v2:
Extended Multi-Turn Depth: 4 to 8 turns per dialogue (mean 5.30 turns).
Never-End-Prematurely: Zero premature goodbyes when staying up late or relaxing.
AI Self-Identity: Rich answers when users ask about the AI's plans, sleep, and… See the full description on the dataset page: https://huggingface.co/datasets/yacdev/somali-100k-saaxiib-conversations-v2.somali-stt-dataset-multi-speaker-v1
Dataset Structure
The dataset contains the following columns:
text: The Somali sentence (transcription).
audio: The audio file sampled at 24,000 Hz.
speaker_id: Unique integer ID (1 to 11) representing each of the 11 speakers.
Metadata & Search Keywords
Language: Somali (so)
Speakers: 11 unique voices (balanced gender representation)
Audio Quality: 24kHz, mono, clean audio
Total Rows: 1,200
Total Duration: ~1.66 Hours (99.86 Minutes)
Intended Use: Fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/laki35/somali-stt-dataset-multi-speaker-v1.robotsim-soma-source-shards-validated-v01SomaSphere_3D_Modelsanv-data-ke-somali
