datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semanticspray-plusplus
Dataset Card for SemanticSpray++ Multimodal (MCAP)
A FiftyOne build of the SemanticSpray++ dataset (Piroli, Dallabetta, Kopp,
Walessa, Meissner & Dietmayer; Institute of Measurement, Control, and
Microtechnology, Ulm University, with BMW AG), a multimodal labeled dataset
for testing camera, LiDAR, and radar perception in wet-surface "vehicle
spray" conditions. This build repackages the 36-scene labeled subset
(SemanticSpray++'s own contribution on top of the earlier… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/semanticspray-plusplus.semantic-segmentation-test-sampleThis dataset contains 10 examples of the segments/sidewalk-semantic dataset (i.e. 10 images with corresponding ground-truth segmentation maps).
semantic-scholarSemantic_Segmantation_Datasetsbilingual-rlhf-financial-semantics
🧠 RLHF Semantic Contribution Dataset: Financial & Policy Bilingual Corpus
Dataset name: sunwang4gptplus/bilingual-rlhf-financial-semanticsCreated by: Sun WangLanguages: English, Mandarin ChineseLicense: MITTags: RLHF, bilingual, semantic_contribution, ux_tone, human_feedback, financial_policy, openai_auditTasks: text-classification, text-generation, reinforcement-learning
📦 Dataset Overview
This dataset was created by Sun Wang as part of a multi-day Reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/sunwang4gptplus/bilingual-rlhf-financial-semantics.SemanticSTF
📌 SemanticSTF Dataset
SemanticSTF is a real multimodal LiDAR dataset collected under adverse weather conditions including rain, snow, and fog, for autonomous driving research.
It provides synchronized LiDAR point clouds, RGB images, and per-point semantic labels of 20 classes, designed for 3D semantic segmentation and sensor fusion tasks.
The dataset contains train/val/test splits, camera intrinsics/extrinsics, and high-quality annotations aligned at the frame level.… See the full description on the dataset page: https://huggingface.co/datasets/AR-X/SemanticSTF.PhenoBench_images_semanticsSemantic-Search-dataset-for-EE5327701modal-semantics-reasoning
Modal Semantics Reasoning
Can a language model change its answer when the rules of modal logic change?
Each example contains the same premises and conclusion under two semantic
specifications. Only one rule about possible worlds or objects changes, and
the correct answer changes with it. Automated theorem provers verify every
label.
This dataset accompanies Same Formulas, Different Semantics: Do Language
Models Follow Modal Logic Specifications?
Dataset subsets… See the full description on the dataset page: https://huggingface.co/datasets/sileod/modal-semantics-reasoning.Semantic-SVG-Benchmark
Semantic SVG Benchmark
A benchmark of 203 SVG files annotated with human-written semantic object-decomposition trees:
every rendered shape (<path>, <rect>, <circle>, …) in each SVG is assigned to a named semantic
object (e.g. judge, gavel), and objects may be further decomposed into parts
(e.g. Bamboo planter → pot, bamboo). It is the evaluation benchmark of
Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing
(EMNLP 2026). The annotations are ours; the SVGs… See the full description on the dataset page: https://huggingface.co/datasets/KU-MIIL/Semantic-SVG-Benchmark.semantic_search_qualityThis is a synthetic dataset containing documents from Wikipedia, Cosmopedia, and CNN DailyMail news datasets containing search keywords that you can use for checking the quality of your semantic search engine
We have also released a library for generating your own synthetic dataset on your own data so you can perform tests.
To build your own datasets, refer Semantic Synth
SemanticScholarCSFullTextWithOpenAlexTopicsSemanticSeg
Dataset Card for SemanticSeg
This semantic segmentation dataset introduced in the paper Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation.
This dataset is used to train the segmenter.
Dataset Details
Dataset Description
SemanticSeg contains around 16 segmentation categories, with each category containing at least 2k instances. The varying cut rates across categories can also help the segmenter learn… See the full description on the dataset page: https://huggingface.co/datasets/Syon-Li/SemanticSeg.semantic-space-artifacts
Semantic Space artifacts
Contents of data/processed/ for Semantic Space:
memmap-ready embedding matrices, a compressed fastText OOV model, HistWords
decade files, a 200k-row 3D layout, and published bias word lists.
This dataset is the runtime artifact store. A Space or local backend with
HF_ARTIFACTS_REPO set downloads it on first boot when manifest.json is missing.
Licenses
Dual-licensed at the file level. The Hub card license is PDDL because
GloVe and HistWords… See the full description on the dataset page: https://huggingface.co/datasets/rexheng/semantic-space-artifacts.semantic-scholar-scraper
Semantic Scholar Scraper · Papers, Authors, Citations & Venues
Scrape academic research papers, authors, citations, venues, and open-access metadata from Semantic Scholar API. Features rate-limit backoff resilience and pay-per-event pricing.
Rows in this dataset
450
Fields
19
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/semantic-scholar-scraper.repro-stable-simulation-ready-tabletop-layout-generation-semantics-physics-dual-system-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Semantic-Scholar-PapersSemantic-Search-V1-14K
Dataset Card for "Semantic-Search-V1-14K"
More Information needed
semantic_sentence_similarity_ES
Dataset Card for "semantic_sentence_similarity_ES"
This dataset is based on https://huggingface.co/datasets/PlanTL-GOB-ES/sts-es, which includes the datasets presented at the SemEval 2014 and 2015 shared tasks on sentence similarity (see the link for more info about the citations). It also includes data from SemEval 2017.
semantic-scholar-manajemen-proyek
Semantic Scholar: Manajemen Proyek
Batch Fetched from Semantic Scholar API, with parameters:
query = '"manajemen proyek" | "project management" | "metodologi manajemen proyek" | "teknik manajemen proyek" | "alat manajemen proyek" | "keterampilan manajemen proyek" | "tantangan manajemen proyek" | "risiko manajemen proyek" | "studi kasus manajemen proyek" | "tren manajemen proyek" | "manajemen proyek di berbagai industri" | "manajemen proyek dalam organisasi" | "manajemen proyek… See the full description on the dataset page: https://huggingface.co/datasets/derhan/semantic-scholar-manajemen-proyek.semanticsegmentationandposeestimationfromrgbdRGB-D dataset for instance segmentation (from RGB or depth) and pose estimation of individual objects. Data has been generated by randomizing bin contents in Webots.
Each instance contains a mask image as well meta data containing labels, position, and size of each object.
You can create your own data by opening webots_grasp.wbt in the world directory using Webots.
semantic_seg_ATL
Dataset Card for "semantic_seg_ATL"
More Information needed
Semantic_similarity_deduplicated_reasoning_data_english
Semantic_similarity_deduplicated_reasoning_data_english
数据集描述
Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.semantic-song-embeddingsws-semantics-simnrel
Dataset Card for WS353-semantics-sim-and-rel with ~2K entries.
Dataset Summary
License: Apache-2.0. Contains CSV of a list of word1, word2, their connection score, type of connection and language.
Original Datasets are available here:
https://leviants.com/multilingual-simlex999-and-wordsim353/
Paper of original Dataset:
https://arxiv.org/pdf/1508.00106v5.pdf
semantic_searchVLM_semantics_SLO_benchmark
VLM Semantics SLO Benchmark
VLM Semantics SLO is a Slovenian multimodal benchmark for studying cultural and semiotic reasoning in vision-language models. It goes beyond object recognition by asking models to interpret visual hierarchy, spatial relations, colour and mood, composition, cultural symbols, metaphor, denotation and connotation, intertextuality, communicative intent, and relevance to Slovenia.
The released JSON contains 4,950 image-level records. Every record has ten… See the full description on the dataset page: https://huggingface.co/datasets/maticmatusek/VLM_semantics_SLO_benchmark.semanticSearchSemantic_Segmentation_CE
Dataset Card for "Semantic_Segmentation_CE"
More Information needed
semantic-scholar-index
