datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flachwitze-de
Artificial-Unintelligence: Deutsche Flachwitze
Ein kuratierter NLP-Datensatz mit deutschen Flachwitzen, Kalauern und Dad Jokes.
Jeder Witz ist strukturiert in Setup, Punchline, Wortspiel-Erklärung und einen Cringe-Score (1-5).
📊 Statistik
Gesamtanzahl: 7749 Witze
Durchschnittlicher Cringe-Score: 3.83 / 5.0
Kategorien:
Wortwitz: 7749
📋 Schema
Spalte
Typ
Beschreibung
id
string
Eindeutiger Identifier (AU-DE-XXXXXX)
setup
string… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Unintelligence/flachwitze-de.DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence
DoD Public Affairs Use of Artificial Intelligence
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 150 document-grounded question-and-answer records based on DoD Instruction 5400.19, “Public Affairs Use of Artificial Intelligence,” effective July 28, 2025.
The source establishes Department of Defense policy, responsibilities, and procedures for the appropriate use of artificial-intelligence capabilities in… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5400-19-Public-Affairs-Use-of-Artificial-Intelligence.brazilian-math-physics-qa
Brazilian Math & Physics QA
English | Português do Brasil
English
Summary
Brazilian Portuguese question-answer pairs covering mathematics, physics, chemistry, and related educational subjects. Each record contains a user question and an assistant answer in chat/SFT format.
Examples: 19.082
Train: 18.148
Validation: 934
Language: Brazilian Portuguese (pt-BR)
Schema
{"id":"qa_...","subject":"fisica","category":"mecanica-geral"… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/brazilian-math-physics-qa.TaskTrove
TaskTrove
TaskTrove is an open-source collection of agentic task datasets, released by the OpenThoughts-Agent team. It is the task complement to AgentTrove — the agent traces in AgentTrove were generated by running models against these task datasets using the Harbor framework.
v3.2 (current) — replaced the old swegym task dataset (laion__swegym-tasks-patched-validated-v2, 989 tasks) with laion/swegym-tasks-patched-validated-v5 (2,438 tasks, patched + validated). No other… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Production-Units/TaskTrove.AgentTrove
AgentTrove
AgentTrove is the largest open-source collection of agentic interaction traces to date, released by the OpenThoughts-Agent team. It contains 1,696,847 rows drawn from 219 source datasets spanning code repair, shell scripting, mathematical problem-solving, competitive programming, and general computer-use tasks.
At 1.7 million rows, AgentTrove is 4× the size of the Nemotron Terminal Corpus (430 K rows), the previous largest open-source agentic trace dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Production-Units/AgentTrove.agent-data-collection
Dataset Card for OpenHands Agent Logs
This dataset consists of multi-turn dialogues between a simulated human and an LLM-based agent interacting in a virtual operating system environment. Each conversation involves the agent reasoning about and solving command-line tasks through execute_bash and related actions.
Dataset Format
Each file in the dataset is a .json file, structured as a list of instances. Each instance contains:
id: A unique identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Production-Units/agent-data-collection.
