datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
basic-skillsinstruct-data-basics-smollm-H4Datasets of basic instructions and answers for SmolLM-Instruct models trainings: it includes answers to greetings and questions such as "Who are you". This dataset was included in training of SmolLM-Instruct v0.2 but we didn't notice that it had an impact on model generations.
We recommend using this generic larger dataset of multi-turn everyday conversations: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k
BasicSpatialAbility
[ACL'25 Main] Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics
[!IMPORTANT]
You can find the sample testing code on GitHub!
This dataset is a benchmark designed for evaluating Multimodal Large Language Models' Basic Spatial Abilities based on authentic Psychometric theories. It is structured specifically to support both Zero-shot and Few-shot evaluation protocols.
Split Name
Role
Description
test
Query Set… See the full description on the dataset page: https://huggingface.co/datasets/EmbodiedCity/BasicSpatialAbility.cs336-basics-collection
CS336 Assignment 1 — Pre-tokenized Data & BPE Tokenizers
This repository contains preprocessing artifacts produced for Stanford CS336: Language Modeling from Scratch, Spring 2025 — Assignment 1: Basics.
It includes:
pre-tokenized TinyStories train/validation data,
pre-tokenized OpenWebText (OWT sample) train/validation data,
byte-level BPE vocabularies and merge tables for both datasets.
The main purpose of this repository is to avoid repeating the relatively expensive… See the full description on the dataset page: https://huggingface.co/datasets/victorhu493/cs336-basics-collection.gstest6basicsvmBasicSR_SR_testbasic_sentence_transforms
Dataset Card for Active/Passive/Logical Transforms
Dataset Summary
This dataset is a synthetic dataset containing structure-to-structure transformation tasks between
English sentences in 3 forms: active, passive, and logical. The dataset also includes several
tree-transformation diagnostic/warm-up tasks.
Supported Tasks and Leaderboards
[TBD]
Languages
All data is in English.
Dataset Structure
The dataset consists of several subsets, or… See the full description on the dataset page: https://huggingface.co/datasets/rfernand/basic_sentence_transforms.basic-skillsbasic_shapes_object_detection
Basic Shapes Object Detection
Description
This Basic Shapes Object Detection dataset has been created to test fine-tuning of object detection models. Fine-tuning some model to detect the basic shapes should be rather easy: just a bit of training should be enough to get the model to do correct object detection quite fast.
Each entry in the dataset has a RGB PNG image with a white background and 3 basic geometric shapes:
A blue square
A red circle
A green triangle
All… See the full description on the dataset page: https://huggingface.co/datasets/driesverachtert/basic_shapes_object_detection.tas-basics-seed-001
README — TAS Expertise Dataset
Overview
This dataset is a seed corpus for fine-tuning LLMs on Task-Agnostic Step (TAS) decomposition.Each example demonstrates how to break down a high-level goal into structured TAS entries with:
Inputs → what is needed
Process → the step itself
Outputs → expected result
Definition of Done (DoD) → criteria for completion
The dataset is intended as a foundation for teaching models systematic task decomposition across design… See the full description on the dataset page: https://huggingface.co/datasets/dokii/tas-basics-seed-001.cooking-knowledge-basics
Comprehensive Cooking Knowledge Q&A Dataset
This dataset (cooking_knowledge.csv) contains a rich collection of synthetically generated Question-Answer (Q&A) pairs covering diverse aspects of cooking knowledge, with particular emphasis on food chemistry, flavor pairing, cooking techniques, dietary accommodations, and culinary traditions. The data was created using a large language model with advanced reasoning capabilities, prompted with various grounded contexts and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ktiyab/cooking-knowledge-basics.xlerobot-final-v1-xlerobot_mobile_basics_h264_v1This dataset was created using LeRobot.
Dataset Description
XLeRobot Final v1 whole-body teleoperation smoke test
A public, real-hardware integration record for an XLeRobot-compatible 0.4.0
two-wheel configuration. This is an engineering smoke-test dataset, not a
household-task training benchmark.
Recorded system
Two Seeed Studio Pro SO-101 arms (leader/follower teleoperation)
Differential two-wheel base
Pan/tilt head
Three synchronized RGB… See the full description on the dataset page: https://huggingface.co/datasets/NOJIMA21/xlerobot-final-v1-xlerobot_mobile_basics_h264_v1.Prompting-Basics-Kita
📝 Prompting Basics für Erzieher: Die Kamera-Regel
Die Qualität einer KI-generierten Bildungs- und Lerngeschichte hängt massiv von den Rohdaten ab, die du ihr übergibst. KI-Systeme brauchen neutrale, präzise Fakten, um pädagogisch wertvoll arbeiten zu können – und absoluten Schutz der kindlichen Privatsphäre.
Dieser Guide zeigt, wie Beobachtungen optimal für KI-Systeme aufbereitet werden.
1. Die eiserne Datenschutz-Regel: Anonymisierung
Namen ersetzen: Nutze… See the full description on the dataset page: https://huggingface.co/datasets/Earlychildhoodeducation/Prompting-Basics-Kita.mycppython-basics-50-docenglish-environmental-science-basics-30ohada_basicsDebian_Linux_BasicsDebian Linux Basic!
SO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoreSO-Python_basics_QA-filtered-2023-tanh_scoreSO dataset of python tag data and "Python basics and Envirinment" subcategory
Question filters:
images
links
code blocks
Q_Score > 0
Answer_count > 0
Answers filters:
images
links
code blocks
Scores are tanh applied to scaled with AbsMaxScaler to IQR range of Original SO Answers' scores
SO_Python_basics_QA_human_prefContrastive dataset for Stack Overflow python basics QA with augmentations:
SO-SO comparisons: 6166
Par-SO comparisons: 0
SO-Par comparisons: 36366
Gen-SO comparisons: 0
SO-Gen comparisons: 87114
Gen-Par comparisons: 0
Par-Gen comparisons: 0
Gen-Gen comparisons: 0
Par-Par comparisons: 55494
Paraphrasing model: humarin/chatgpt_paraphraser_on_T5_base
english-business-ethics-basics-30english-software-engineering-basics-30english-project-management-basics-30english-supply-chain-basics-30pt_basicsphonetically diverse standalone words, letters, diphtongs and basic greetings
english-risk-management-basics-30english-leadership-skills-basics-30english-human-resources-basics-30
