datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UAVReason_depth
🗺️ UAVReason Depth
Depth maps and depth statistics for UAV-native multimodal reasoning and generation
📌 News
Paper: Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
arXiv: arXiv:2604.05377
VQA / caption / generation annotations: jarvissun/UAVReason_vqa
This dataset is released as part of UAVReason, introduced in the paper above.Please cite the paper if you use this dataset.
🧭 Overview
UAVReason Depth… See the full description on the dataset page: https://huggingface.co/datasets/jarvissun/UAVReason_depth.MMArt-PPR10k
MMArt-PPR10k Dataset
The MMArt-PPR10k is a multimodal dataset specifically created for research into the instruction-driven agentic image retouching task. It is built upon the original PPR10k dataset and offers rich, paired image data, user instructions, and information on the Lua/XMP tools used in Lightroom.
Dataset Structure
The dataset is organized into a hierarchical folder structure. Each data sample corresponds to a specific image pair and its related files, located… See the full description on the dataset page: https://huggingface.co/datasets/JarvisArt/MMArt-PPR10k.jarvis-hud-assetslm-eval-results-shyamieee-JARVIS-v2.0-private
Dataset Card for Evaluation run of shyamieee/JARVIS-v2.0
Dataset automatically created during the evaluation run of model shyamieee/JARVIS-v2.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-JARVIS-v2.0-private.UAVReason_vqa
🚁 UAVReason: VQA / Caption / Generation Annotations
A UAV-native benchmark for aerial visual reasoning, captioning, temporal understanding, and cross-modal generation
📌 News
Paper: Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
arXiv: arXiv:2604.05377
Depth data: jarvissun/UAVReason_depth
This dataset is released as part of UAVReason, introduced in the paper above.Please cite the paper if you use this dataset.… See the full description on the dataset page: https://huggingface.co/datasets/jarvissun/UAVReason_vqa.stark-nano-java-corpusopen-schematics
Open Schematics Dataset
A comprehensive dataset of electronic schematics from hardware projects. This dataset is designed for training AI models on circuit design, component recognition, and hardware engineering tasks.
Dataset Description
This dataset contains electronic schematic files along with their visual representations, component information, and metadata from various hardware projects.
Dataset Structure
Each record in the dataset contains:
schematic:… See the full description on the dataset page: https://huggingface.co/datasets/jarvisemitra/open-schematics.Medical-ASR-ENjarvis_voiceOpenMascotAI_Jarvis
OpenMascotAI Jarvis — 페르소나 없는 자비스(도구 호출) 데이터셋
OpenMascotAI 데스크톱 마스코트의 자비스 모드(마스코트가 윈도우를
실제로 조작하는 모드)를 순정 LLM에 가르치기 위한 한국어 function-calling 데이터셋입니다. 캐릭터 페르소나 데이터셋
(MelissaJ/ProjectLucia_Hera의 tools config)에서 캐릭터 요소를 전부 걷어낸 공개 참고용 판입니다.
config
split
행 수
tools
train / eval
5,494 / 322
A Korean function-calling dataset for the "Jarvis mode" of OpenMascotAI (a desktop mascot that actually operates
Windows). It is the persona-free edition of the character dataset: user… See the full description on the dataset page: https://huggingface.co/datasets/MelissaJ/OpenMascotAI_Jarvis.ovos-wake-word-bench-picovoice-jarvis
OVOS wake_word bench — picovoice-jarvis
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-jarvis.ovos-wake-word-bench-synthetic-wakewords-hey_jarvis
OVOS wake_word bench — synthetic-wakewords-hey_jarvis
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
OpenVoiceOS/synthetic-wakewords.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-synthetic-wakewords-hey_jarvis.jarvis-shapes
JARVIS Shapes
Validated 3D design specs/scripts for JARVIS Chat's design engine. Public so the app can fetch the latest shapes at launch.
design/design_memory_seed.jsonl — the seed JARVIS loads ({prompt, script}).
design/all.jsonl — full merged dataset.
design/spec_shapes.jsonl — the 250 spec-generated geometry shapes.
jarvisham10000-skin-lesions
HAM10000 Skin Lesion Dataset
Dataset Description
The HAM10000 dataset contains 10,015 dermatoscopic images of pigmented skin lesions.
Classes
akiec: Actinic Keratoses and Intraepithelial Carcinoma
bcc: Basal Cell Carcinoma
bkl: Benign Keratosis-like Lesions
df: Dermatofibroma
mel: Melanoma
nv: Melanocytic Nevi
vasc: Vascular Lesions
Usage
from datasets import load_dataset
dataset = load_dataset("jarvisit/HAM10000-skin-lesions")
jarvis-XRD-cleanjarvis-voice-samplesjarvis-assistantsynthetic-wakeword-hey_jarvis
synthetic-wakeword-hey_jarvis
Synthetic wake-word audio for training and benchmarking OVOS wake-word
plugins, covering the phrase "hey jarvis".
Every clip is machine-generated text-to-speech. No human recording is
included, and no natural voice is reproduced. Machine-generated audio carries
no copyright of its own, so this dataset is published CC-BY-4.0 and is free
to use, redistribute and build on, including for model training.
Produced with support from the NGI0 Commons Fund.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/synthetic-wakeword-hey_jarvis.Jarvis_Piper_TTS_VoiceJARVIS_DFT_3D_8_18_2021
Cite this dataset Choudhary, K., Garrity, K. F., Reid, A. C. E., DeCost, B., Biacchi, A. J., Walker, A. R. H., Trautt, Z., Hattrick-Simpers, J., Kusne, A. G., Centrone, A., Davydov, A., Jiang, J., Pachter, R., Cheon, G., Reed, E., Agrawal, A., Qian, X., Sharma, V., Zhuang, H., Kalinin, S. V., Sumpter, B. G., Pilania, G., Acar, P., Mandal, S., Haule, K., Vanderbilt, D., Rabe, K., and Tavazza, F. JARVIS DFT 3D 8 18 2021. ColabFit, 2023. https://doi.org/10.60732/a9dd64f6… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_DFT_3D_8_18_2021.details_shyamieee__JARVIS-v3.0jarvis_deep_dataset
JARVIS Deep Dataset
A large-scale instruction-tuning dataset for persona fine-tuning, built by rewriting 70,000 general-purpose instruction-response pairs from WizardLM's Evol-Instruct dataset into the voice of JARVIS — Tony Stark's AI assistant from the Iron Man films.
Unlike narrow persona datasets that only cover a handful of topics, this dataset preserves the full breadth and complexity of WizardLM's instructions — coding, reasoning, writing, analysis, mathematics, and… See the full description on the dataset page: https://huggingface.co/datasets/roger33303/jarvis_deep_dataset.JARVIS_DFT_3D_12_12_2022
Cite this dataset Choudhary, K., Garrity, K. F., Reid, A. C. E., DeCost, B., Biacchi, A. J., Walker, A. R. H., Trautt, Z., Hattrick-Simpers, J., Kusne, A. G., Centrone, A., Davydov, A., Jiang, J., Pachter, R., Cheon, G., Reed, E., Agrawal, A., Qian, X., Sharma, V., Zhuang, H., Kalinin, S. V., Sumpter, B. G., Pilania, G., Acar, P., Mandal, S., Haule, K., Vanderbilt, D., Rabe, K., and Tavazza, F. JARVIS DFT 3D 12 12 2022. ColabFit, 2023. https://doi.org/10.60732/e9e65ccd… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_DFT_3D_12_12_2022.R1-Onevision-with-Systemjarvis-synth-textdetails_VAIBHAV22334455__JARVIS
Dataset Card for Evaluation run of VAIBHAV22334455/JARVIS
Dataset automatically created during the evaluation run of model VAIBHAV22334455/JARVIS on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_VAIBHAV22334455__JARVIS.jarvis-piper-checkpointsJARVIS_OQMD_no_CFID
Cite this dataset Kirklin, S., Saal, J. E., Meredig, B., Thompson, A., Doak, J. W., Aykol, M., Rühl, S., and Wolverton, C. JARVIS OQMD no CFID. ColabFit, 2023. https://doi.org/10.60732/82cb32aa
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_untxw8nljf92_0
Visit the ColabFit Exchange to search additional datasets by author, description… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/JARVIS_OQMD_no_CFID.resource
