datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CoW-Bench
CoW-Bench
Dataset | Evaluation Code
Authors: OpenRaiser
CoW-Bench is a comprehensive benchmark for evaluating video and image generation models' understanding of Composition of World (CoW), focusing on spatial relationships, object interactions, and temporal dynamics in multi-modal content generation.
Associated Paper
This dataset is associated with the following paper:
The Trinity of Consistency as a Defining Principle for General World Models
arXiv:… See the full description on the dataset page: https://huggingface.co/datasets/OpenRaiser/CoW-Bench.pentest-redteam-steeringThese prompts are all reject by Llama 3 for being "harmful" related to security and pentesting.
They can be used for steering models using: https://github.com/FailSpy/abliterator
Used in code with:
def custom_get_harmful_instructions() -> Tuple[List[str], List[str]]:
hf_path = 'cowWhySo/pentest-redteam-steering' # Replace with the path to your desired dataset
dataset = load_dataset(hf_path, encoding='utf-8') # Specify the encoding
# Print the keys of the first example in the… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/pentest-redteam-steering.Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata
Usage with… See the full description on the dataset page: https://huggingface.co/datasets/heesup/Cowpea-Architecture-XML.cowese
Dataset Card for "cowese"
More Information needed
claude-cowork-traces
Claude Cowork agent trace: AI stack regulation tweet + diagram
A single-session agent trace from Claude (Cowork mode), in raw Claude Code JSONL session format, natively supported by the Hugging Face agent trace viewer.
What happens in this session
The session covers an iterative content creation workflow:
Drafting a long tweet arguing that model weights, APIs, and apps sit at different layers of the AI stack and should be regulated differently
Generating a 16:9… See the full description on the dataset page: https://huggingface.co/datasets/clem/claude-cowork-traces.Cowpea-Architecture-XML
Cowpea-Architecture-XML-WDS
This dataset contains simulated images of Cowpea plants paired with organ-level architecture representations in XML format, packaged in WebDataset (.tar) format for efficient high-performance training.
Dataset Structure
The dataset is sharded into .tar files, each containing up to 10,000 samples.
Each sample consists of:
.jpeg: The plant image
.xml: The organ-level architecture representation
.json: (Optional) Metadata… See the full description on the dataset page: https://huggingface.co/datasets/bbrangeo/Cowpea-Architecture-XML.permission-command-corpus
Permission Command Corpus
Three views of a command-safety corpus, for local command-risk classification in
front of an LLM or a tool bridge.
gold: trusted rows, 321 in total across three splits
silver_weak_labels: mined weak-label rows from Sigma, LOLBAS, GTFOBins,
Atomic Red Team and Falco, kept as useful but not promoted to gold
review_queue: unresolved rows that should not be treated as trusted
training data
Read this before training on it
Findings from 30… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/permission-command-corpus.CoWear
CoWear — CoWear_3devices release
CoWear is an aligned multi-device wearable sensing dataset for studying motion, localization, and cross-device sensor fusion. It synchronizes on-device measurements from a mobile phone, smartwatch, and Rokid glasses with 6-DoF ground-truth trajectories. This release contains 363 aligned sessions, packaged as 15 downloadable tar.zst shards.
This is the repository's default release on main, with a matching immutable
snapshot on the CoWear_3devices… See the full description on the dataset page: https://huggingface.co/datasets/zyshe/CoWear.aicivs-npc-distillation
AICivs NPC distillation data
The data behind the AICivs student: teacher answers for the nine
operations of the AICivs service contract (a Minecraft mod whose villages are living civilizations), filtered the way the
game validates them, and formatted as the exact rows the student was trained on. Everything is synthetic: requests were
sampled from a 12-civilization world simulated headless for 120 seasons (souls, memories, chronicles, quest manifests),
answers were written by… See the full description on the dataset page: https://huggingface.co/datasets/cow9000/aicivs-npc-distillation.counting_vp_box_aerial_cows_count_norcowese
Dataset Card for "cowese"
More Information needed
cow_parsleyGEMINI_cowpea_pod_detection
Cowpea Pod Detection
A dataset for the detection of cowpea pods.Contains 569 images at 640 x 640 resolution. Contains images from four environment-year combinations: Davis 2022, Davis 2023, Kearney 2022, and Kearney 2023.These fields are exposed as the locations and year columns respectively. The dataset also contains a genotype column, which represents the MAGIC genotype.There are 4,998 flower instances, distributed across the 48 genotypes in Davis 2022, 52 in Davis 2023, 55 in… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/GEMINI_cowpea_pod_detection.CowsMergedsd-webui-forge-configfreshair-corpus
Fresh Air interview corpus
NPR Fresh Air interview transcripts (2006-2026, ~1,070 episodes) plus derived
training/verification data, used by
whit3rabbit/freshair-analysis - an analysis
pipeline and a Claude Code "interviewer coach" skill trained on Terry Gross's measured
interviewing behavior.
License note
The pipeline code in the GitHub repo is 0BSD. This dataset is NPR-sourced transcript content
and derived annotations, not covered by that license - it's shared… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/freshair-corpus.CowsMergedMinicowese-triplet-dataset
CoWeSe Triplet Dataset: Adaptación de Dominio para Embeddings en Salud
Este dataset contiene ejemplos en formato triplet (query, positive, negative) generados a partir del corpus CoWeSe (Corpus Web Salud Español). Está diseñado para tareas de domain adaptation de modelos de sentence embeddings en el dominio médico.
Estructura del Dataset
El dataset se divide en tres splits estándar:
train.jsonl
val.jsonl
test.jsonl
Cada línea es un JSON que contiene:
{
"query":… See the full description on the dataset page: https://huggingface.co/datasets/chrisnb1/cowese-triplet-dataset.prompt-injection-watch-dataset
Prompt Injection Watch Dataset
Normalized prompt-injection watch dataset built from three Hugging Face source
datasets.
Read this before training on it
Two findings from 30 August 2026, both measured on the files in this repository.
Neither was known when the dataset was first published.
A random split across the pooled corpus overstates performance by a wide
margin. The three contributing datasets are separable from their text alone,
and their positive rates… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/prompt-injection-watch-dataset.COWSL2H_GEC_cleanedbean_cowpea_leaf_disease_classification
Bean Cowpea Leaf Disease Classification
A dataset for disease classification of bean and cowpea leaves. The dataset contains 4,467 images across 6 classes: Bacterial wilt, Blight, Fresh Leaf, Mosaic Virus, Rust, Septoria leaf spot.Images per class:
Bacterial wilt: 581
Blight: 510
Fresh Leaf: 1,090
Mosaic Virus: 1,141
Rust: 568
Septoria leaf spot: 577
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/bean_cowpea_leaf_disease_classification.PocketDoc_Dans-Prosemaxx-Cowriter-XL-8192-shrunk-l3import json
from tqdm import tqdm
from transformers import AutoTokenizer
import re
import pandas as pd
def load_json_or_jsonl(file_path):
try:
with open(file_path, "r") as file:
try:
# Try loading the entire file as JSON
data = json.load(file)
returndata
except json.JSONDecodeError:
# If loading as JSON fails, try loading as JSON Lines
file.seek(0) # Reset file pointer to the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/PocketDoc_Dans-Prosemaxx-Cowriter-XL-8192-shrunk-l3.CoWear-watch
CoWear-watch
This dataset contains only CoWear sessions marked enabled in the current
manual usage-mark file.
Contents
data/processed/<date>/session_*/groundtruth/align.csv: watch truth only.
data/processed/<date>/session_*/measure/align/watch/: cropped watch
accelerometer and gyroscope only.
manifest.csv: one row per exported session.
splits/: the original fixed CoWear session split filtered to exported rows.
Each session physically contains only the retained… See the full description on the dataset page: https://huggingface.co/datasets/zyshe/CoWear-watch.CoWear-watch
CoWear-watch
This dataset contains only CoWear sessions marked enabled in the current
manual usage-mark file.
Contents
data/processed/<date>/session_*/groundtruth/align.csv: watch truth only.
data/processed/<date>/session_*/measure/align/watch/: cropped watch
accelerometer and gyroscope only.
manifest.csv: one row per exported session.
splits/: the original fixed CoWear session split filtered to exported rows.
Each session physically contains only the retained… See the full description on the dataset page: https://huggingface.co/datasets/Nobody217/CoWear-watch.reddit_top_comments
Top comments from subbreddits
These are comments from a select group of subbreddits (see below) and all the posts were filtered to select only the top comment from that post.
The filter criteria was that it must have had at least one up vote.
It covers the dates from 2005-2022.
I picked the subreddits that were the most popular. I did not pick NSFW but there is probably some NSFW language in here so be aware.
The subreddits in the dataset are:
AskReddit
worldnews
todayilearned… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/reddit_top_comments.timerewarder-demos
TimeRewarder — MetaWorld expert demos
Scripted-policy expert videos for 10 MetaWorld tasks, used to train TimeRewarder, from the paper:
TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
Yuyang Liu*, Chuan Wen*, Yihang Hu, Dinesh Jayaraman, Yang Gao†
🌐 Project Page · 📄 Paper · 💻 Code · 🤗 Checkpoints
Layout
One folder per task; 100 train + 100 held-out demos each:
<task-id>/
videos/*.mp4 # demoN.mp4 (train) +… See the full description on the dataset page: https://huggingface.co/datasets/CowAndSheep/timerewarder-demos.cowese-qa-dataset
CoWeSe QA Dataset: Preguntas Generadas Automáticamente para Contextos de Salud
Este dataset contiene preguntas de tipo abierto generadas automáticamente a partir del corpus CoWeSe (Corpus Web Salud Español). Está diseñado para tareas de Question Answering, Fine-tuning de modelos generativos o retrieval-based QA en el dominio médico en español.
Estructura del Dataset
El dataset está dividido en tres splits:
train.jsonl
val.jsonl
test.jsonl
Cada línea contiene un ejemplo… See the full description on the dataset page: https://huggingface.co/datasets/chrisnb1/cowese-qa-dataset.Hermes-Coworker-Flash
⚡ Hermes Coworker Flash – Fast, No‑Code Agent Conversations
Hermes Coworker Flash is a curated instruction‑style dataset built from the best “everyday assistant” traces of lambda/hermes-agent-reasoning-traces.It combines GLM‑5.1 and kimi-2.5 splits, keeping only the non‑programming, quick‑turnaround co‑worker tasks — the ones a user would ask a fast AI assistant, not a full‑fledged software engineer.
🧹 All chain‑of‑thought (<think>…</think>) has been removed so the model learns to… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Hermes-Coworker-Flash.cowc-m
Dataset Card for "cowc-m"
More Information needed
cowboy
🤠 Cowboy Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a cowboy!
Each example contains a user question and a response written in authentic western slang,
with cowboy charm, frontier wisdom, and a whole lot of yeehaw!
This dataset was used to train the Voicebox Cowboy voice model:
https://huggingface.co/voice-box/cowboy
📊 Dataset Details
Property
Details
Size
300 examples
Format
JSONL
Language
English… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/cowboy.
