willow
Datasets
All datasets matching “willow”gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/WillowVoiceAI/gaming-500-hours.willow-surface-code-detection-events
Willow Surface-Code Detection Events (ingested)
Detection events and logical-observable flips derived from Google's Willow
below-threshold surface-code dataset (Zenodo
10.5281/zenodo.13273331), rotated
surface code at distances 3, 5, 7 in X and Z memory.
Each row is one experimental shot. Detection events are derived from the raw
device measurement records with Stim's measurement-to-detector converter, using
the per-shot sweep bits, and validated to reproduce the dataset's… See the full description on the dataset page: https://huggingface.co/datasets/ShayManor/willow-surface-code-detection-events.Nemotron-Pretraining-Code-v3
Nemotron-Pretraining-Code-v3
Dataset Description:
The Nemotron-Pretraining-Code-v3 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset is intended to improve the coding capabilities of LLMs.
The Nemotron-Pretraining-Code-v3 dataset contains the metadata corresponding to the raw source-code update to our Nemotron-Pretraining-Code-v2 and Nemotron-Pretraining-Code-v1… See the full description on the dataset page: https://huggingface.co/datasets/WillowVoiceAI/Nemotron-Pretraining-Code-v3.WillowNLtoFOL
Dataset Card for WillowNLtoFOL
Dataset Summary
WillowNLtoFOL is a high-quality, structurally diverse dataset designed for evaluating the compositional generalization capabilities of neural semantic parsing models on the Natural Language to First-Order Logic (NL-to-FOL) translation task.
Translating natural language into formal logic requires nuanced linguistic understanding and the systematic recombination of learned structures. Existing datasets often suffer… See the full description on the dataset page: https://huggingface.co/datasets/iedeveci/WillowNLtoFOL.willowodia-indextts2-processed
