datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BigWorld-PoC
BigWorld PoC
One persistent, 16-day synthetic workplace ecosystem with two competing enterprises, one government agency, three consumers, six employees and six dedicated native Hermes/bubblewrap computers. Twelve distinct pinned synthetic Persona 8B records condition actual MiroFish/OASIS actors. Companies change goals and offers in response to competition, outcomes, policy and fictional geopolitical changes. Employee files, conversations and commitments persist between… See the full description on the dataset page: https://huggingface.co/datasets/ojus1/BigWorld-PoC.pocketgull-nih-who-clinical-dpo
📚 PocketGull NIH & WHO Clinical Preference DPO Dataset
Organization: PocketGull LLC (Oregon SOS: 258869891)Curator: Phillip Gear (CMS NPI: 1487569752 | ORCID: 0009-0008-1372-5381)License: Creative Commons Attribution 4.0 International (CC-BY-4.0)Open Science DOI: 10.5281/zenodo.20647514
📌 Dataset Summary
Gold-standard Direct Preference Optimization (DPO) chosen vs rejected pairs grounded in NIH MedQuAD, WHO mhGAP guidelines, and ClinicalTrials.gov protocols… See the full description on the dataset page: https://huggingface.co/datasets/philgear/pocketgull-nih-who-clinical-dpo.Dans-SystemmaxxThis dataset would not be possible without Gryphe and Caitlyns amazing work on https://huggingface.co/datasets/Gryphe/Sonnet3.5-SlimOrcaDedupCleaned and https://huggingface.co/datasets/cgato/SlimOrcaDedupCleaned
ksl-pose-dictionary-poc
KSL Pose Dictionary (PoC)
한국수어(KSL) text-to-pose 시제품용 keypoint 데이터셋.
docent_AI_sign_research_02 프로젝트에서 생성. Neural Sign Actors (CVPR 2024) 접근법을 KSL에 적용하는 Path B (Dictionary-based) 시제품의 핵심 데이터셋.
개요
자산
갯수
키포인트
sldict keypoint (국립국어원 한국수어사전)
1,444 단어
OpenPose 137 (RTMW-DW-L-M 추출)
NIASL2021 gloss segmentation keypoint (재난 안전 도메인)
2,287 base gloss
OpenPose 137 (NIASL 원본)
Hybrid sign index
4,511 unique signs
단어 → keypoint 경로 매핑
Stage 1 학습 corpus
20,085 samples… See the full description on the dataset page: https://huggingface.co/datasets/Trotquonalize/ksl-pose-dictionary-poc.pocket-mechanic-distilled
Pocket Mechanic: distillation dataset
4,365 (sensor window → diagnostic explanation) pairs for fine-tuning small language models to do OBD-II car-fault diagnosis with calibrated repair-cost estimates and shop-upsell awareness.
Used to train MindFreakGamer/gemma-4-E2B-pocket-mechanic. The student reaches 81.3% of teacher quality on a blind A/B judge (n=100).
Submission to the Hugging Face Build Small Hackathon, Backyard AI track (June 2026). Code:… See the full description on the dataset page: https://huggingface.co/datasets/MindFreakGamer/pocket-mechanic-distilled.vim-command-pocket
NickIBrody/vim-command-pocket
Vim Command Pocket Dataset is a small offline seed built from official Vim help pages.
It is intentionally lightweight and is meant to serve as the starting point for a larger Vim corpus.
Source pages
https://vimhelp.org/usr_02.txt.html
https://vimhelp.org/quickref.txt.html
Split sizes
{
"total_examples": 747,
"train": 599,
"validation": 74,
"test": 74,
"sources": [
"https://vimhelp.org/quickref.txt.html"… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/vim-command-pocket.word-mean-nice-llm-sft-poc
Fine-tuning Dataset
This dataset was generated for fine-tuning language models with deterministic
training-format transforms.
Dataset Details
Session ID: session_10761044
Repository: monostate/word-mean-nice-llm-sft-poc
Number of Examples: 191
Format: JSONL (JSON Lines)
Training Type: sft
Chat Template: chatml
Render Strategy: both
Generator Model: kimi-k2.5
Generated: 2026-03-16T09:59:01.806970
Dataset Structure
Each example contains either structured… See the full description on the dataset page: https://huggingface.co/datasets/monostate/word-mean-nice-llm-sft-poc.
