datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pi-coding-sessionssankakuany4hdmi-g1-real-picoovos-wake-word-bench-picovoice-computer
OVOS wake_word bench — picovoice-computer
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-computer.pi-computer-use-sessions
Coding agent session traces for thomasmustier/pi-computer-use-sessions
This dataset contains redacted coding agent session traces collected while working on https://github.com/tmustier/pi-computer-use. The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction, secret scanning, visual review where applicable, and LLM review.
Source git repo: https://github.com/tmustier/pi-computer-use
Data… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-computer-use-sessions.ovos-wake-word-bench-picovoice-alexa
OVOS wake_word bench — picovoice-alexa
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-alexa.ovos-wake-word-bench-picovoice-smart-mirror
OVOS wake_word bench — picovoice-smart-mirror
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.ovos-wake-word-bench-picovoice-jarvis
OVOS wake_word bench — picovoice-jarvis
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-jarvis.ovos-wake-word-bench-picovoice-snowboy
OVOS wake_word bench — picovoice-snowboy
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-snowboy.ovos-wake-word-bench-picovoice-view-glass
OVOS wake_word bench — picovoice-view-glass
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-view-glass.sankaku_jsonRWKV-notebook-assets
RWKV notebook assets
Various asset files, used in RWKV notebook tutorial examples and demos
PicoAudioThe dataset utilized in PicoAudio
pico_bigbio_processeddanbooru_jsonguanaco-pico-100-samplesThis dataset is a subset of the Open Assistant dataset, which you can find here: https://huggingface.co/datasets/OpenAssistant/oasst1/tree/main
This subset of the data only contains 100 samples of the highest-rated paths in the conversation tree.
This dataset was used to train Guanaco with QLoRA.
For further information, please see the original dataset.
License: Apache 2.0
PiCo-DATA-general
language:
en license: openrail task_categories:
text-generation tags:
instruction-tuning
identity
safety
alignment
reasoning
pico-1b size_categories:
50K<n<100K
PiCo-1B Instruction Dataset
A high-quality, high-variability instruction dataset for PiCo-1B (ArcOffical/PiCo-1B), a ~1.46B parameter dense decoder-only transformer trained from scratch by the Arc Develop Team.
Purpose
This dataset teaches PiCo-1B who it is, what it can do, what it should do, what it… See the full description on the dataset page: https://huggingface.co/datasets/ArcOffical/PiCo-DATA-general.PiCo-dataset-general-7b
PiCo-7B Instruction Dataset
A high-quality, high-variability instruction dataset for PiCo-7B (ArcOffical/PiCo-7B), an approximately 6.95B-parameter (~7B) large language model featuring an Adaptive Hierarchical Mixture of Experts (AHMoE) architecture. The official model card describes 30 layers, including 15 MoE layers and 15 dense layers, with approximately 2.63B active parameters per token and a 131,072-token context window.[^1]
Purpose
This dataset teaches… See the full description on the dataset page: https://huggingface.co/datasets/ArcOffical/PiCo-dataset-general-7b.Liquid_V1_7B-pico-aurora-vidgen-multiturn-annotationslung_Proton_therapy_pico20241211-pico
