datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepLearningarxiv_deep_learning_python_research_code_functions_summaries
Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries
Dataset Summary
AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.arxiv_deep_learning_python_research_code
ArXiv Deep Learning Python Research Code
A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.
Dataset Summary
Statistic
Value
Total files
391,496
Total size
1.49 GB
Source repos
34,099
Time span
ArXiv inception through July 2023
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks
GPUMemNet and GPUUtilNet Dataset
This dataset accompanies the paper
“GPU Memory and Utilization Estimation for Training-Aware Resource
Management: Opportunities and Limitations.”
It contains synthetic deep learning training configurations and their measured
GPU memory consumption and utilization characteristics.
Dataset configurations
The dataset is divided into separate configurations because MLP, CNN, and
Transformer workloads use different feature schemas.… See the full description on the dataset page: https://huggingface.co/datasets/ehyo/GPU-Resources-Estimation-for-Deep-Learning-Training-Tasks.arxiv_small_nougat
Dataset Description
The "arxiv_small_nougat" dataset is a collection of 108 recent papers sourced from arXiv, focusing on topics related to Large Language Models (LLM) and Transformers. These papers have been meticulously processed and parsed using Meta's Nougat model, which is specifically designed to retain the integrity of complex elements such as tables and mathematical equations.
Data Format
The dataset contains the parsed content of the selected papers, with special… See the full description on the dataset page: https://huggingface.co/datasets/deep-learning-analytics/arxiv_small_nougat.Ko.HelpSteer원본 데이터셋: nvidia/HelpSteer
ko.SHP
🚢 Korean Stanford Human Preferences Dataset (Ko.SHP)
이 데이터셋은 자체 구축한 번역기를 활용하여 stanfordnlp/SHP 데이터셋을 번역한 것입니다.
아래의 내용은 해당 번역기로 README 파일을 번역한 것입니다. 참고 부탁드립니다.
If you mention this dataset in a paper, please cite the paper: Understanding Dataset Difficulty with V-Usable Information (ICML 2022).
Summary
SHP는 요리에서 법률 조언에 이르기까지 18가지 다른 주제 영역의 질문/지침에 대한 응답에 대한 385K 집단 인간 선호도 데이터 세트이다.
기본 설정은 다른 응답에 대 한 한 응답의 유용성을 반영 하기 위한 것이며 RLHF 보상 모델 및 NLG 평가 모델 (예: SteamSHP)을 훈련 하는 데… See the full description on the dataset page: https://huggingface.co/datasets/nlp-with-deeplearning/ko.SHP.cdg-AICourse-Level3-DeepLearning
Learner & EnfuseBot: Exploring the role of Regularization in Neural Network Training - Generated by Conversation Dataset Generator
This dataset was generated using the Conversation Dataset Generator script available at https://cahlen.github.io/conversation-dataset-generator/.
Generation Parameters
Number of Conversations Requested: 500
Number of Conversations Successfully Generated: 500
Total Turns: 6655
Model ID: meta-llama/Meta-Llama-3-8B-Instruct
Generation Mode:… See the full description on the dataset page: https://huggingface.co/datasets/cahlen/cdg-AICourse-Level3-DeepLearning.IITM_Intro_to_Deep_Learning_Nppe1_exam_dataset
🧠 IITM Intro to Deep Learning & GenAI NPPE1 — Age & Gender Prediction Dataset
This dataset was prepared for the IIT Madras "Intro to Deep Learning & GenAI NPPE1" competition hosted on Kaggle.It contains face images and metadata used for multi-task learning — predicting both age (regression) and gender (classification) from image inputs.
📦 Dataset Structure
Files Included
File
Description
train/
Folder containing training face images.… See the full description on the dataset page: https://huggingface.co/datasets/AyusmanSamasi/IITM_Intro_to_Deep_Learning_Nppe1_exam_dataset.DeepLearningCommonFactor_DLvsIPCADeepLearningAndCommonFactorsDeeplearning-CommonFactors
