datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantaai-crypto_assets
semantaai-crypto_assets
Semanta AI market dataset in unified raw + gold + labels standard.
updated at UTC: 2026-07-22T22:01:35.864528+00:00
target end UTC: 2026-07-22T20:20:00+00:00
raw rows: 4284625
gold rows: 6174008
Layers:
raw/: canonical 5m exchange/vendor bars with quality flags.
gold/: cleaned, synchronized, TA/regime/context features. Past/present only.
labels/: future-derived path-dependent labels. Do not join into features before train/test split design.
semantaai-fx-majors
semantaai-fx-majors
Semanta FX dataset with two layers: raw and gold.
raw rows: 32531270
gold rows: 46788944
symbols: 28
end UTC: 2026-06-30T23:55:00+00:00
source: Dukascopy public historical candles
semasia-mnist
Latents for mnist (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on mnist, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with datasets and convert to… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-mnist.semantaai-fx-majors-legacy-20260403
semantaai-fx-majors
Semanta AI dataset with two layers: raw and gold.
raw rows: 32032242
gold rows: 46072167
from floor UTC: 2006-01-01T00:00:00+00:00
end UTC: 2026-04-03
semantic-memorization-partial-2023-09-03This dataset is a partial computation of metrics (memorized token frequencies, non-memorized token frequencies, sequence frequencies) needed for research.
semasia-oxford-flowers
Latents for oxford-flowers (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on oxford-flowers, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-oxford-flowers.pythia-semantic-memorization-perplexities
Dataset Card for "pythia-semantic-memorization-perplexities"
More Information needed
semasia-cifar100
Latents for cifar100 (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on cifar100, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with datasets and… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-cifar100.semasia-imagenet-1k
Latents for imagenet-1k (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on imagenet-1k, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with datasets… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-imagenet-1k.semasia-fashion_mnist
Latents for fashion_mnist (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on fashion_mnist, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-fashion_mnist.semantaai-fx-majors-gold5m-legacy-gamma
semantaai-fx-majors-gold5m
Semanta AI majors gold 5m layer extracted from fx-majors.
semantic-ecology
NuBerea Semantic Ecology
Feature-indexed analysis layer for the study of canon formation, part of the
NuBerea corpus estate of biblical and patristic texts. It describes the
"semantic ecology" of ancient biblical and related literature — how semantic
domains are used across texts, how candidate texts were received over time,
and which measurable features accompany canonical inclusion — packaged as a
set of ready-to-load configurations.
Attribution
NuBerea project.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/semantic-ecology.semantaai-fx-other-gold5m
semantaai-fx-other-gold5m
Semanta AI FX other gold 5m layer extracted from fx-other.
semantaai-fx-other
semantaai-fx-other
Semanta FX dataset with two layers: raw and gold.
raw rows: 30653399
gold rows: 13232988
symbols: 33
end UTC: 2026-06-30T23:55:00+00:00
source: Dukascopy public historical candles
SemanticVLA-TraceX-240K-DROID
SemanticVLA TraceX 240K · DROID
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The DROID component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of DROID · Franka · Open-X-Embodiment DROID… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-DROID.semanta-crypto-assets
semantaai-crypto_assets
Semanta AI market dataset in unified raw + gold standard.
updated at UTC: 2026-05-17T11:39:09.295802+00:00
target end UTC: 2026-05-17T23:55:00+00:00
raw rows: 4182508
gold rows: 6026532
SemanticPotentialRoutingTelemetry
Semantic Potential Routing Telemetry
Version 2.0 — a systematic, packet-level benchmark of training-free potential-field routing against
classical routing under stochastic congestion, dynamic topologies and microburst traffic.
Every episode is a network simulation in which routing is a physical field: each flow's destination is the
grounded, attractive well of a discrete Poisson equation on the graph Laplacian, congested buffers inject
repulsive current, and packets follow the… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/SemanticPotentialRoutingTelemetry.pile-semantic-memorization-filter-results
Dataset Card for "pile-semantic-memorization-filter-results"
More Information needed
semantic-duplicatesMaCoCu-sl-tokenized
Dataset Card for MaCoCu-sl Multi-Tokenized
Dataset Description:
This dataset provides a pre-tokenized version of the Slovene web corpus MaCoCu.
It includes the original text data and metadata from MaCoCu-sl, augmented with token IDs and token counts generated by several
popular large language model tokenizers. The goal is to facilitate research and experimentation by providing ready-to-use tokenized data,
saving computational resources during repeated setups.
Licensing and… See the full description on the dataset page: https://huggingface.co/datasets/SemantikaEU/MaCoCu-sl-tokenized.semantaai-fx-other-legacy-20260403
semantaai-fx-other
Semanta AI FX other dataset:
raw: 5m OHLCV
gold: 15m, 1h, 4h, 1d
Gold 5m is published separately in Grencape/semantaai-fx-other-gold5m.
EEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics.
The raw EEG data and the datasheet are available at https://osf.io/xh3g5/.
See code repository for benchmark results.
EEG data acquisition:
Explanations of the variables:
event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.dexmg_dataset_semantic_goalsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "robomimic",
"total_episodes": 1014,
"total_frames": 326707,
"total_tasks": 1,
"total_videos": 3042,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1014"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ishika/dexmg_dataset_semantic_goals.generation-semantic-memorization-filterssemasia-tiny-imagenet
Latents for tiny-imagenet (timm)
This repository hosts precomputed latent representations (embeddings) extracted from timm image-classification backbones on tiny-imagenet, released as part of SEMASIA — a large-scale resource for studying semantic communication, cross-model latent space alignment, and explainability.
Each config corresponds to a single model;
only that model's Parquet files are read on load_dataset.
Usage
Load with… See the full description on the dataset page: https://huggingface.co/datasets/spaicom-lab/semasia-tiny-imagenet.memories-semantic-memorization-filter-results
Dataset Card for "memories-semantic-memorization-filter-results"
More Information needed
sec-10k-semantic-dedupedsemantaai-fx-majors-gold5m
semantaai-fx-majors-gold5m
Semanta AI majors gold 5m layer extracted from fx-majors.
generation-semantic-filtersgeneration-semantic-intermediate-filters
