datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deep_learning_2025_vision_jointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy",
"total_episodes": 50,
"total_frames": 10128,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025_vision_joint.ROCOv2-X-Ray-radiology_-cycle-1
ROCOv2 X-Ray, Report Generation Pilot (Cycle 1)
A 28-row pilot testing whether short ROCO captions can be expanded into report-shaped training pairs.
Why this exists
ROCOv2 captions are one or two clipped sentences, often written to make a teaching point rather than to read as a radiological finding. A vision-language model trained directly on them learns to produce clipped captions, not reports.
This pilot tested a different approach: take the caption and its… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/ROCOv2-X-Ray-radiology_-cycle-1.ROCOv2-X-Ray-radiology
ROCOv2 X-Ray Subset
Radiographs extracted from ROCOv2 (Radiology Objects in COntext, version 2), for training vision-language models on X-ray interpretation.
What this is
ROCOv2 spans many imaging modalities. This subset keeps only the X-ray studies, so that a model can be trained on a single modality rather than learning across CT, MRI, ultrasound and radiography at once.
Rows
4,254
Split
train
Size
~977 MB
Modality
X-ray only… See the full description on the dataset page: https://huggingface.co/datasets/deepLEARNING786/ROCOv2-X-Ray-radiology.deep_learning_2025This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy",
"total_episodes": 50,
"total_frames": 10128,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025.radimagenet-vqa-500-test
🩺 RadImageNet VQA 500 Test Subset
This dataset contains a curated 500-example test audit subset for Medical Visual Question Answering based on RadImageNet.
📊 Dataset Summary
Total Samples: 500 test VQA pairs
Modalities: CT, MRI, X-ray (Abdomen, Brain, Chest/Lung, Ankle/Foot, Hip, Knee)
Question Types: Open-ended & Closed (Yes/No)
Organization: VQA-DeepLearning
💻 Usage
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/VQA-DeepLearning/radimagenet-vqa-500-test.actresses
Dataset Card for "actresses"
More Information needed
Nvidia-DeepLearningExamplesCode from https://github.com/NVIDIA/DeepLearningExamples
INFO: Found 4341 text files - 2024-Jan-27_02-13
INFO: Train size: 4123
Validation size: 109
Test size: 109
deep_learning_books_dataset
Deep Learning Books Dataset
Dataset Information
Features:
page_no: Integer (int64) - Page number in the book.
page_content: String - Text content of the page.
Splits:
train: Training split.
Number of examples: 474
Number of bytes: 1,030,431
Download Size: 509,839 bytes
Dataset Size: 1,030,431 bytes
Dataset Application
This dataset "deep_learning_books_dataset" contains text data from various pages of books related to deep learning.
It can be… See the full description on the dataset page: https://huggingface.co/datasets/Falah/deep_learning_books_dataset.deep-learning-domain-qavqa-rad
Dataset Card for VQA-RAD
Dataset Description
VQA-RAD is a dataset of question-answer pairs on radiology images. The dataset is intended to be used for training and testing
Medical Visual Question Answering (VQA) systems. The dataset includes both open-ended questions and binary "yes/no" questions.
The dataset is built from MedPix, which is a free open-access online database of medical images.
The question-answer pairs were manually generated by a team of… See the full description on the dataset page: https://huggingface.co/datasets/VQA-DeepLearning/vqa-rad.cdg-AICourse-Level3-DeepLearning
Learner & EnfuseBot: Exploring the role of Regularization in Neural Network Training - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 500
Number of Conversations Successfully Generated: 500
Total Turns: 6655
Model ID: meta-llama/Meta-Llama-3-8B-Instruct
Generation Mode:… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-AICourse-Level3-DeepLearning.dataset3deep_learning_books
Dataset Card for "deep_learning_books"
More Information needed
dataset2dataset1Laila
Dataset Card for "Laila"
More Information needed
dataset51Lailalyrics_genre_dataset_mediumdeeplearning_lmmLaila_New
Dataset Card for "Laila_New"
More Information needed
dataset4fall2025-deeplearning-noisy-pairs-newfall2025-deeplearning-noisy-pairs-new-2fall2025-deeplearning-noisy-100kfall2025-deeplearning-noisy-test-10fall2025-deeplearning-noisy-pairs-new-colab-1.7
