datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sscc-compact-av
SSCC compact balanced multimodal subset
This private derived dataset contains 288 synchronized SSCC clips from 15 medium-load,
clean operating conditions at speeds 60, 80, and 100. It retains recorder FLAC audio,
four anti-aliased 25 kHz vibration channels in compressed float32 NPZ, and five sparse frames
from both iOS and Android videos. Five sample IDs retain both unchanged source MP4s for
presentation and loader tests.
The subset is balanced between normal and fault states… See the full description on the dataset page: https://huggingface.co/datasets/DesanSilva/sscc-compact-av.b1k-joint-v5-compactlemonseed-compact-foundation-cogen
lemonseed-compact-foundation-cogen
LemonSeed — compact foundation teacher-co-gen training (v2).
Contents
intelligent_compact_foundation_train_v2.jsonl (88 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
interstate-licensure-compact-participation
Interstate Professional Licensure Compact Participation by State
Canonical, always-current version: https://referencesource.org/interstate-licensure-compact-participation/
Machine-readable: https://referencesource.org/interstate-licensure-compact-participation/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-17
Stale after: 2026-11-15 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 406
Which U.S.… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/interstate-licensure-compact-participation.PersonalFinance-CoTR-V2-Compact
Personal Finance Reasoning V2 - Compact Edition
This document describes a condensed version of the Personal Finance Reasoning dataset, specifically adapted for training smaller language models (e.g., 1B-4B parameters).
1. Introduction & Motivation
The landscape of financial AI benchmarks is currently dominated by applications in corporate finance, algorithmic trading, and general financial knowledge extraction. While valuable, these benchmarks often overlook the critical… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/PersonalFinance-CoTR-V2-Compact.CompactDS-102GB-queriesobsidian-bases-query-v2-compactkanitakorn-qwen-v27-compact-hard-clean-20260614
Qwen v27 Compact Hard ThaiExam Mix
Compact fallback mix for Kanitakorn. It raises fresh hard ThaiExam coverage while limiting replay dilution. Run generated-record audit, SFT inspection, and contamination screening before GPU training.
support-triage-rlvr-compactloomstack-compact-expert
