datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StreetviewLLM_DataStreetVision-10K
StreetVision-10K
Each sample contains:
A system prompt instructing the model to act as an OSINT/geospatial expert
A user message with a street-level photo and the instruction to determine coordinates
An assistant response with ground-truth coordinates in <direct_lon_lat_output>longitude,latitude</direct_lon_lat_output> format
Format
Each line is a JSON array of ChatML messages:
[
{"role": "system", "content": "..."},
{"role": "user", "content": [… See the full description on the dataset page: https://huggingface.co/datasets/mishl/StreetVision-10K.maib-incident-reports-5K
MAIB Incident Type Dataset
The MAIB Incident Type Dataset contains short textual descriptions of marine accidents and incidents reported by the UK Marine Accident Investigation Branch (MAIB).Each record includes a short narrative and a corresponding incident-type label (e.g. Grounding / Stranding, Fire / Explosion, Collision).This dataset enables research and experimentation in maritime safety text classification and domain-specific NLP.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/baker-street/maib-incident-reports-5K.StreetMathDatasetsf-streets-gpt-audio-resultsHarley_Street_Institute_Aesthetic_FAQ
Harley Street Institute Aesthetic FAQ — Nepali Q&A Dataset
1. Overview
This dataset is a Nepali-language (Devanagari script) collection of question–answer pairs in ShareGPT format, covering frequently asked questions about aesthetic medicine (एस्थेटिक मेडिसिन) — specifically aimed at doctors, nurses, dentists, and other healthcare professionals considering or building a career in medical aesthetics (botox, dermal fillers, injectable treatments, aesthetic training… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Harley_Street_Institute_Aesthetic_FAQ.StreetMath
Street Math Approximation Dataset
A dataset for training language models on mental math approximation and reasoning skills.
Dataset Description
This dataset contains mental math problems designed to teach approximation strategies and reasoning. Each example includes:
Input: A mental math question requiring approximation
Output: The approximate answer using mental math techniques
Exact Answer: The precise mathematical result
Bounds: Acceptable approximation range (±10%… See the full description on the dataset page: https://huggingface.co/datasets/Chiung-Yi/StreetMath.r_streetwear-gemini-2.0-flash-thinking-exp-1219-CustomShareGPTstreet-scene-descriptions
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
street_scene_descriptions
This dataset contains over 10,000 street-level image samples accompanied by detailed textual scene descriptions and rich metadata including geolocation, time, weather, and infrastructure details. Each entry features a natural language caption describing urban and residential environments captured from vehicle perspectives under varying lighting and… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/street-scene-descriptions.Wall-Street-Agent-Tracessepitori-street-v18th-street-latinasStreetMath
