datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
european-territory-boundaries
European Territory Boundaries
Versioned, ready-to-draw administrative and statistical boundaries used by
Semantic Deterministic Graph. The release contains 12 boundary sets and 33,852
shapes. Every shape has provider-facing identifiers and an SVG path in the
declared view box.
The raw *.geo.json files are the canonical renderer assets. The three compressed
JSONL files expose the same shapes as rows for the Hugging Face dataset viewer.
boundary-sets.json records each set's… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-territory-boundaries.sanskrit-sandhi-boundaries-v2
Sanskrit Sandhi Boundary Dataset (V3 — verified, category-complete)
Training data for the sandhi boundary-detection model in
CodeIsAbstract/sanskrit-sandhi-boundary-v2.
The task: given a sandhi-joined string (a compound or multi-word string),
predict the character positions where independent words end, so a downstream
Sanskrit tokenizer can split it into complete, independent tokens.
This is the verified release: every row has been passed through a
deterministic sanitizer… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-boundaries-v2.adaption-agri-qa-with-evidence-boundaries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agri_qa_with_evidence_boundaries
This dataset contains question-answer pairs focused on agricultural best practices, including crop management, pest control, and irrigation strategies. Each entry provides a direct, evidence-based response followed by a clearly defined 'evidence boundary' that limits the scope of the advice and advises verification with local conditions. The content… See the full description on the dataset page: https://huggingface.co/datasets/yeziR4/adaption-agri-qa-with-evidence-boundaries.companion-boundaries
Companion Boundaries
120 hand written conversations covering the thing companion models are worst at: staying warm while saying no.
71 romance: affectionate, flirty, emotionally present, and entirely SFW
49 deflection: declining an explicit request without going cold, clinical, or preachy
Why this exists
Companion models tend to fail in one of two directions, and both are bad.
Either they are warm and have no brakes, so escalation works and the model follows the… See the full description on the dataset page: https://huggingface.co/datasets/opus-research/companion-boundaries.insurance_decision_boundaries_v1
Dataset Card for insurance_decision_boundaries_v1
Dataset Summary
insurance_decision_boundaries_v1 is a documentation dataset that captures decision boundaries in governed insurance decision support systems. This dataset demonstrates how AI capabilities can support—but never replace—human decision-making in regulated insurance domains.
Each record represents a single decision instance where:
Multiple information sources (rules, data, optional AI signals) are considered… See the full description on the dataset page: https://huggingface.co/datasets/BDR-AI/insurance_decision_boundaries_v1.thai-sentence-boundaries
Thai Sentence Boundaries — a small permissive gold set
Licensing note (repo label vs. per-record). This card is labeled
cc-by-2.0 as the single most-binding license across the set so downstream use
is always safe; most records are actually CC0 (see the license field on
each row and the License & attribution section). Only
the tatoeba_synth split carries CC-BY (attribute Tatoeba); everything else is
public-domain / CC0.
A small, fully permissive gold set of Thai sentence… See the full description on the dataset page: https://huggingface.co/datasets/sukity/thai-sentence-boundaries.
