datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ParaDLC-Bench
ParaDLC-Bench
ParaDLC-Bench (Parallel Detailed Localized Captioning Benchmark) is a benchmark for multi-region localized captioning that jointly evaluates caption quality and inference efficiency. It extends DLC-Bench from single-region evaluation to concurrent multi-region evaluation, explicitly stressing a model's ability to describe many regions at once without cross-region interference.
📄 Paper |
💻 Code |
🤖 PerceptionDLM
Key… See the full description on the dataset page: https://huggingface.co/datasets/MSALab/ParaDLC-Bench.clr_motion_planning_hw_7agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
marine_vla_dataset
Marine VLA Dataset
Vision-Language-Action dataset for autonomous marine vessel navigation using SmolVLA.
Dataset Structure
LeRobot-style format with 37 episodes, 12175 frames:
data/
episode_000000/
episode_data.json # frame-by-frame labels + metadata
observation.images.camera_0/
000000.jpg # 640x480 RGB frames
000001.jpg
...
episode_000001/
...
dataset_info.json # schema, label names, stats… See the full description on the dataset page: https://huggingface.co/datasets/MSaalaamaa/marine_vla_dataset.clr_mujoco_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omy",
"total_episodes": 2,
"total_frames": 1161,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/msavchen-nasa/clr_mujoco_dataset.clr_motion_planning_hw_8xrays-and-gradcamCogneo-Crypto-MSAfunction-graphs-trclr_motion_planning_datasetclr_motion_planning_hw_6splitter-dataset
Splitter Dataset
Bu dataset, matematiksel grafik sorularını adım adım çözüm yaklaşımıyla işlemek için tasarlanmıştır. Her görsel için, soruyu mantıksal alt sorulara (sub-questions) ayıran bir "Question Decomposer" sistemi geliştirilmiştir.
🎯 Amaç
Model, herhangi bir matematik grafiği sorusunu çözmeye çalışmadan önce, soruyu bir bütün olarak ele alıp mantıklı ve anlamlı adımlara ayırır. Bu yaklaşım, karmaşık soruların daha kolay anlaşılmasını ve çözülmesini sağlar.… See the full description on the dataset page: https://huggingface.co/datasets/Msalcann/splitter-dataset.clr_motion_planning_hw_v4clr_motion_planning_v3clr_motion_planning_hw_v1clr_motion_planning_3clr_motion_planning_hw_4test0msa_datasetsixteenth_MSA_renderedtwitter-s1co0jq-2024.11.10-1855554234135244828-msal36YW8qPxH7Ay-part1clr_motion_planning_hw_2clr_motion_planning_hw_3grafikLife_Style_Slepp
Sleep Health and Lifestyle Analysis
Course: Introduction to Data Science (IDS F24)Instructor: Dr. M Nadeem MajeedProject Deadline: December 10, 2025
Project Overview
This project performs comprehensive analysis on the Sleep Health and Lifestyle Dataset to understand relationships between sleep patterns, lifestyle factors, and health outcomes. The project includes:
15+ EDA analyses (summary stats, distributions, correlations, boxplots, pairplots, skewness, outliers… See the full description on the dataset page: https://huggingface.co/datasets/MSAMI1506/Life_Style_Slepp.
