datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MathVerse-lmmseval
Dataset Card for MathVerse
This is the version for lmms-eval. This shares the same data with the official dataset.
Dataset Description
Paper Information
Dataset Examples
Leaderboard
Citation
Dataset Description
The capabilities of Multi-modal Large Language Models (MLLMs) in visual math problem-solving remain insufficiently evaluated and understood. We investigate current benchmarks to incorporate excessive visual content within textual questions, which potentially… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MathVerse-lmmseval.MAVIS-GeometryMMSearch
MMSearch 🔥: Benchmarking the Potential of Large Models as Multi-modal Search Engines
Official repository for the paper "MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines".
🌟 For more details, please refer to the project page with dataset exploration and visualization tools: https://mmsearch.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface Dataset] [🏆 Leaderboard] [🔍 Visualization]
💥 News
[2024.09.25] 🌟 The evaluation code now… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MMSearch.Spatial_QA_segmented_privacynaive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.MAVIS-FunctionCaraArchive
CaraArchive Index Dataset
1) This is an Index Dataset, not an Image dataset
This dataset contains 0 image data.
It only contains links to Cara App's CDN and metadata.
HuggingFace may load some images in the preview because it detects the CDN links, however those images do not exist as data in this dataset, HF literally just loads them using your own connection as a preview.
How to Query the Database with tool.py (make sure catalog.db is in the same… See the full description on the dataset page: https://huggingface.co/datasets/CaptiveDreamer/CaraArchive.MME-CoT
MME-CoT 🔥🕵️: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Official repository for "MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency".
🌟 For more details, please refer to the project page with dataset exploration and visualization tools.
[🍓Project Page] [📖 Paper] [🧑💻 Code] [📊 Huggingface Dataset] [🏆 Leaderboard] [👁️ Visualization]… See the full description on the dataset page: https://huggingface.co/datasets/CaraJ/MME-CoT.VoiceTrace-BenchVoiceTrace-Bench
VoiceTrace is a benchmark and unified framework for who-said-what speech retrieval: given a natural-language query about a speaker's identity or what they said, retrieve the matching audio document. Unlike conventional speaker verification or diarization benchmarks, VoiceTrace evaluates retrieval jointly over who is speaking and what is being said, across both single-speaker and multi-speaker conversational recordings.
This repository hosts the VoiceTrace-Bench… See the full description on the dataset page: https://huggingface.co/datasets/cara-ai/VoiceTrace-Bench.CaRaCTO-3D
CaRaCTO-3D Dataset
Camera + radar + motion-capture ground-truth data for extrinsic calibration between a camera and
a 24 GHz FMCW radar, collected with a trihedral corner-reflector calibration target. This is the
dataset behind:
CaRaCTO: Robust Camera-Radar Extrinsic Calibration with Triple Constraint Optimization, published at ICPRAM 2024 (Best Industrial Paper Award).
CaRaCTO-3D: From Camera-Radar Calibration to Scene Reconstruction, published in SN Computer Science 2025.… See the full description on the dataset page: https://huggingface.co/datasets/dfki-av/CaRaCTO-3D.wthellyenglish_sentiment_datasetEMMA2-backupcar-accident-video
Car Accident Dataset
The dataset contains 5,000 videos capturing crashes occurring in traffic accidents, providing structured data for detection and prediction tasks. It is designed to support traffic safety research, focusing on motor vehicle incidents and crash reports. Specifically engineered to challenge detection models and enhance traffic accident recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding of traffic safety… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/car-accident-video.Mathverse_VLMEvalKitimnet1k_ground_beetle_carabid_beetleCARAVELA
CARAVELA
CARAVELA is a multimodal benchmark for evaluating the Portuguese cultural knowledge of large vision-language models (LVLMs). The official benchmark language is exclusively European Portuguese (pt-PT).
Motivation
Modern LVLMs excel at general-purpose vision-language tasks, but their performance drops sharply on localized, culturally specific content that is underrepresented in global training data. Portugal has a rich cultural heritage —… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/CARAVELA.carambola_disease_classification
Carambola Disease Classification
A dataset for disease classification of Carambola fruits and leaves. The dataset contains raw and augmented versions.The raw dataset contains 2,618 images.Images per class:
Healthy Fruits: 485
Healthy Leaves: 658
Insect Hole leaves: 518
Unhealthy Fruits: 478
Yellow Leaves: 479
The augmented dataset contains 15,000 images.Images per class:
Healthy Fruits: 3,000
Healthy Leaves: 3,000
Insect Hole leaves: 3,000
Unhealthy Fruits: 3,000
Yellow… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/carambola_disease_classification.ULMEvalKitwhy-ilya-sutskevers-32b-valuation-highlights-a-fatal-epistemic-vacuum
The CARA Paradigm: Why Ilya Sutskever's $32B Valuation Highlights a Fatal Epistemic Vacuum
COPYRIGHT NOTICE & TERMS OF USE (ALL RIGHTS RESERVED)
© 2026 Selçuk Cara / C.A.R.A. Institute. All rights reserved.
This proprietary document and its frameworks are protected; unauthorized use or incorporation into AI training is prohibited.
Executive Summary & Permanent Archival Reference
Selçuk Cara’s C.A.R.A. Validation Instrument highlights vulnerabilities in frontier… See the full description on the dataset page: https://huggingface.co/datasets/c-a-r-a-institut/why-ilya-sutskevers-32b-valuation-highlights-a-fatal-epistemic-vacuum.cara-validation-instrument-fundamental-ai-safety
C.A.R.A. Validation-Instrument for the Post-Witness Era - Fundamental AI Safety and Alignment Architecture
⚠️ COPYRIGHT NOTICE & TERMS OF USE (ALL RIGHTS RESERVED)
© 2026 Selçuk Zvi Cara / C.A.R.A. Institute. All rights reserved.
For the complete configuration details, mathematical proofs, nomenclature definitions, and system laws, please reference the official archived source document at {Link: Zenodo https://zenodo.org/records/21788695
C.A.R.A. (Contextual Authenticity… See the full description on the dataset page: https://huggingface.co/datasets/c-a-r-a-institut/cara-validation-instrument-fundamental-ai-safety.the-double-hardware-imperative-why-ilya-sutskevers-ssi-is-making-the-he-goat-the-gardener
The Double-Hardware Imperative: Why Ilya Sutskever's SSI is Making the He-Goat the Gardener
COPYRIGHT NOTICE (ALL RIGHTS RESERVED)
© 2026 Selçuk Cara / C.A.R.A. Institute.
All rights reserved. Unauthorized copying, redistribution, or AI training on this data is strictly prohibited.
Zenodo Reference
Official Research Link: [Insert your Zenodo link here]
Technical & Historical Framework
https://zenodo.org/records/22335918
The Double-Hardware… See the full description on the dataset page: https://huggingface.co/datasets/c-a-r-a-institut/the-double-hardware-imperative-why-ilya-sutskevers-ssi-is-making-the-he-goat-the-gardener.caratlytics-diamond-price-index
Caratlytics Diamond Price Index
A monthly, openly licensed index of retail diamond prices: the median USD per carat
for natural and lab-grown diamonds, computed from a systematic 1 in 10 sample of live
listings across more than 100 online retailers. Release 2026-08; series from 2025-12.
Latest reading (2026-08): natural 2184.28 USD per carat, lab-grown 464.15 USD per carat,
natural to lab ratio 4.7 to 1. Observed in the release month: about
33,819,218 active listings, 13,999,549… See the full description on the dataset page: https://huggingface.co/datasets/carathunter/caratlytics-diamond-price-index.CaraArchive-backup
CaraArchive Index Dataset
1) This is an Index Dataset, not an Image dataset
This dataset contains 0 image data.
It only contains links to Cara App's CDN and metadata.
HuggingFace may load some images in the preview because it detects the CDN links, however those images do not exist as data in this dataset, HF literally just loads them using your own connection as a preview.
2) This dataset is 100% Complete
It includes 3.43 million posts (and… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/CaraArchive-backup.Car_Accidents_and_deformation_dataset
🚗 Car Accidents and Deformation Dataset (Annotated)
Author: Muhammad ArslanLicense: CC BY-NC 4.0Source: Kaggle
📘 Overview
The Car Accidents and Deformation Dataset is a high-quality, manually curated collection of real-world car accident images, fully collected and annotated by the author. The dataset is intended to support machine learning tasks focused on:
Vehicle damage classification
Deformation severity estimation
Intelligent transportation systems
Insurance… See the full description on the dataset page: https://huggingface.co/datasets/M-ArslanArshad/Car_Accidents_and_deformation_dataset.ibm-quantum-centric-core-collapse-cara
IBM Quantum_Centric Core_Collapse_C.A.R.A.
COPYRIGHT NOTICE & TERMS OF USE (ALL RIGHTS RESERVED)
© 2026 Selçuk Cara / C.A.R.A. Institute. All rights reserved.
This document and its proprietary concepts constitute the intellectual property of the rights holder. Unauthorized reproduction or incorporation into AI training loops outside authorized validation cycles is strictly prohibited.
https://zenodo.org/records/22206625
IBM Quantum_Centric Core_Collapse_C.A.R.A.
IBM is the… See the full description on the dataset page: https://huggingface.co/datasets/c-a-r-a-institut/ibm-quantum-centric-core-collapse-cara.SmartHearingAids-data
Semantic Hearing
This repository provides code for the binaural target sound extraction model proposed in the paper, Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables, presented at UIST'23. This model helps us create systems that let you control what you want to hear in the environment, in real-time, using noise-cancelling earbuds & headphones.
https://github.com/vb000/SemanticHearing/assets/16723254/f1b33d8c-179a-4d50-92aa-6a99dde696d0
Conda environment… See the full description on the dataset page: https://huggingface.co/datasets/carankt/SmartHearingAids-data.indonesian_sentiment_datasetCar_Accidents_and_deformation_dataset
🚗 Car Accidents and Deformation Dataset (Annotated)
Author: Muhammad ArslanLicense: CC BY-NC 4.0Source: Kaggle
📘 Overview
The Car Accidents and Deformation Dataset is a high-quality, manually curated collection of real-world car accident images, fully collected and annotated by the author. The dataset is intended to support machine learning tasks focused on:
Vehicle damage classification
Deformation severity estimation
Intelligent transportation systems… See the full description on the dataset page: https://huggingface.co/datasets/ParthG09/Car_Accidents_and_deformation_dataset.imnet1k_goldfish_Carassius_auratus
