datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
grug-moe-mix-swarm
Grug-MoE Data-Mix Experiments
The default config contains the original 840-run Fisher-DSP swarm. The harrier_18t75_d768 config contains the Harrier experiments described below.
Fisher-DSP swarm (default)
840 MoE pretraining runs from the Grug-MoE Fisher-DSP data-mixing swarm (d512, TPU / us-central2).
Each run trains on a distinct data mixture over 168 datakit buckets; the swarm is used to regress
mixture weights → eval loss and predict an optimized pretraining… See the full description on the dataset page: https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm.grundwortschatz-voc-de
WortUniversum German Vocabulary Database
Status: published at
cstr/grundwortschatz-voc-de
(GPL-3.0). The GPL-3.0 licensing is why the app does not bundle this database:
it downloads grundwortschatz-app.db.gz from here on first use. This dataset
is the CC-BY-SA→GPL-3.0 re-distribution form of the data built by the
WortUniversum pipeline. Rebuild
re-push with pipeline/build_hf_datasets.py --upload.
Dataset summary
A curated, enriched lexical database of 10,450… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-de.grundwortschatz-voc-en
WortUniversum English Vocabulary Database
Status: published at
cstr/grundwortschatz-voc-en
(CC-BY-SA 4.0). The shipped app asset is assets/grundwortschatz_en.db.gz in
the WortUniversum / words-universe
repository; this dataset is its CC-BY-SA 4.0 re-distribution form. Rebuild +
re-push with pipeline/build_hf_datasets.py --upload.
Dataset summary
A UK-English lexical database of 11,539 lemmas covering primary-school
vocabulary (CEFR-J A1–B2, YLE… See the full description on the dataset page: https://huggingface.co/datasets/cstr/grundwortschatz-voc-en.rot_gruen_sort_20260724_094754This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/sara2369/rot_gruen_sort_20260724_094754.eval_smolvla-hs-gruppe-03-subtraktiv-1dwesui-grupa-1-neurologia
NeuroSpeechPL
Publiczny eksport HuggingFace zawiera wyłącznie redystrybuowalne audio source=natural. Wiersze TTS są celowo wyłączone z publicznego zbioru danych, ponieważ ich source_license zabrania redystrybucji audio. Pełna lokalna ewaluacja opisana w raporcie korzystała zarówno z nagrań naturalnych, jak i TTS.
Repozytorium zbioru danych HF: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Repozytorium kodu:… See the full description on the dataset page: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia.dwesui-grupa-2-kulinarna
G2-Polish-Culinary-ASR-Evaluation-Corpus
Korpus do ewaluacji systemow ASR jezyka polskiego (domena kulinarna) stworzony
w ramach warsztatow Ewaluacja Systemow Rozpoznawania Mowy (UAM WMI, edycja 2026,
zespol 2). Publikowany podzbior to mowa naturalna z wideo kulinarnych YouTube
(licencja CC-BY) - sluzy do badania odpornosci ASR na szum kuchenny oraz dopasowania
domenowego do specjalistycznego slownictwa (zapozyczenia, miary, liczby).
Pelny eksperyment ewaluacyjny zespolu… See the full description on the dataset page: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna.hs-gruppe-02_filteredThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ned2",
"total_episodes": 14,
"total_frames": 5343,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 21,
"splits": {
"train": "0:14"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AlexanderRoempke/hs-gruppe-02_filtered.hs-gruppe-03gruener_wuerfel_v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 114,
"total_frames": 68219,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:114"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Carlo0001/gruener_wuerfel_v2.hs-gruppe-04CS370-A2-rav4-video-retrieval-Attallah
RAV4 Exterior Video Retrieval Pipeline
This repository contains the data outputs for a custom image-to-video retrieval system. The pipeline processes a target video of a Toyota RAV4, extracts and indexes car part detections frame-by-frame, and clusters those detections into contiguous video clips based on queried objects from the aegean-ai/rav4-exterior-images dataset.
There are two primary data files included in this repository:
1. The Video Index… See the full description on the dataset page: https://huggingface.co/datasets/grugg/CS370-A2-rav4-video-retrieval-Attallah.eval_smolvla-hs-gruppe-04eval_smolvla-hs-gruppe-02amazon-appliances-data-subsetrot_gruen_sorths-gruppe-03-subtraktiv-1hs-gruppe-01-subtraktiv-1zwesui-grupa-5-it-ai
Wykorzystanie ASR do transkrypcji polskich nagrań o tematyce AI
Korpus do ewaluacji systemów ASR języka polskiego stworzony w ramach warsztatów
Ewaluacja Systemów Rozpoznawania Mowy (UAM WMI, edycja 2026, zespół 5).
Zbiór powstał jako część kursu - publikujemy go publicznie, żeby inni badacze
polskiego ASR mogli z niego korzystać i porównywać wyniki na wspólnym benchmarku.
Cel i pytania badawcze
Cel główny:
Porównanie jakości 3 systemów ASR dla spontanicznej… See the full description on the dataset page: https://huggingface.co/datasets/slapekm/zwesui-grupa-5-it-ai.grupo_public_dataamazon-reviews-subset
