sema
Datasets
All datasets matching “sema”semantic-vad-eot
Semantic-VAD EOT
End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible
with livekit/eot-bench-data.
Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered
silence_spans. Per the eot-bench convention the last silence span is the true
end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored).
Splits
For every data type, all shards except the last form the train base; that… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/semantic-vad-eot.SemanticKITTIsemantaai-crypto_assets
semantaai-crypto_assets
Semanta AI market dataset in unified raw + gold + labels standard.
updated at UTC: 2026-07-22T22:01:35.864528+00:00
target end UTC: 2026-07-22T20:20:00+00:00
raw rows: 4284625
gold rows: 6174008
Layers:
raw/: canonical 5m exchange/vendor bars with quality flags.
gold/: cleaned, synchronized, TA/regime/context features. Past/present only.
labels/: future-derived path-dependent labels. Do not join into features before train/test split design.
semantaai-fx-majors
semantaai-fx-majors
Semanta FX dataset with two layers: raw and gold.
raw rows: 32531270
gold rows: 46788944
symbols: 28
end UTC: 2026-06-30T23:55:00+00:00
source: Dukascopy public historical candles
semasia-mnist
Latents for mnist (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on mnist, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with datasets and convert to… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-mnist.FUSU-Fine_grained_Urban_Semantic_Understanding
About:
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:
Bi-temporal high-resolution satellite RGB images with fine-grained annotations.
Monthly revisited Sentinel-2 and Sentinel-1 images.
Details:
1.… See the full description on the dataset page: https://huggingface.co/datasets/sp-juni/FUSU-Fine_grained_Urban_Semantic_Understanding.
