datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv_deep_learning_python_research_code
ArXiv Deep Learning Python Research Code
A curated corpus of Python source code files extracted from GitHub repositories referenced in ArXiv papers. Contains 391,496 files (1.49 GB) filtered to deep learning frameworks, designed for training and evaluating Code LLMs on research-grade code.
Dataset Summary
Statistic
Value
Total files
391,496
Total size
1.49 GB
Source repos
34,099
Time span
ArXiv inception through July 2023
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code.quantum-machine-learning-models
Neura Parse — Quantum Machine Learning Models: Encodings, Kernels, QNNs & Generative/Deep Architectures
A hands-on, code-first vertical on quantum models that learn from data. Spans data encodings/feature maps, variational classifiers, quantum kernels/QSVMs, and quantum neural networks through modern generative and deep architectures (quantum GANs, circuit Born machines, quantum Boltzmann machines, QCNNs, quantum autoencoders, quantum RL, and quantum… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-models.quantum-machine-learning-theory
Neura Parse — Quantum Machine Learning Theory: Trainability, Generalization & Learning From Quantum Data
A research-depth, proof-oriented vertical on the learning theory of quantum models and quantum data. Covers why parameterized quantum circuits train or don't (barren plateaus), what they can represent, when they generalize or provably beat classical models, and — for quantum data — how to predict properties of unknown states/channels with few measurements (classical… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-machine-learning-theory.Fathom-V0.6-Iterative-Curriculum-Learninglearningbench
LearningBench: Scenario-Based Learning for Delivery Leaders
A curated dataset of 361 decision scenarios, 114 caselets, and 9 report samples designed for training and evaluating AI systems on project, programme, and service delivery management skills — now including AI in Delivery Leadership and AIOps scenarios.
Dataset Description
LearningBench provides realistic, workplace-grounded scenarios that test the judgment of project managers, programme managers, and service… See the full description on the dataset page: https://huggingface.co/datasets/shahamitkumar/learningbench.task718_mmmlu_answer_generation_machine_learning
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task718_mmmlu_answer_generation_machine_learning
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task718_mmmlu_answer_generation_machine_learning.smolified-personalized-learning-content-intelligence-platform
🤏 smolified-personalized-learning-content-intelligence-platform
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model taniabiswas232/smolified-personalized-learning-content-intelligence-platform.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 2ad61909)
Records: 260
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset… See the full description on the dataset page: https://huggingface.co/datasets/taniabiswas232/smolified-personalized-learning-content-intelligence-platform.
