boundaries
european-territory-boundaries
European Territory Boundaries
Versioned, ready-to-draw administrative and statistical boundaries used by
Semantic Deterministic Graph. The release contains 12 boundary sets and 33,852
shapes. Every shape has provider-facing identifiers and an SVG path in the
declared view box.
The raw *.geo.json files are the canonical renderer assets. The three compressed
JSONL files expose the same shapes as rows for the Hugging Face dataset viewer.
boundary-sets.json records each set's… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-territory-boundaries.Quantitative_Mapping_of_Computational_Boundaries
Quantitative Mapping of Computational Boundaries
A Statistical Field Theory Approach to Phase Transitions in NP-Hard Problems
Author: Zixi Li (Oz Lee)
Affiliation: Noesis Lab (Independent Research Group)
Contact: lizx93@mail2.sysu.edu.cn
Overview
Classical computability theory tells us that computational boundaries exist (halting problem, P vs NP), but it doesn't answer: where exactly are these boundaries?
This paper presents the first quantitative mapping of… See the full description on the dataset page: https://huggingface.co/datasets/OzTianlu/Quantitative_Mapping_of_Computational_Boundaries.sanskrit-sandhi-boundaries-v2
Sanskrit Sandhi Boundary Dataset (V3 — verified, category-complete)
Training data for the sandhi boundary-detection model in
CodeIsAbstract/sanskrit-sandhi-boundary-v2.
The task: given a sandhi-joined string (a compound or multi-word string),
predict the character positions where independent words end, so a downstream
Sanskrit tokenizer can split it into complete, independent tokens.
This is the verified release: every row has been passed through a
deterministic sanitizer… See the full description on the dataset page: https://huggingface.co/datasets/CodeIsAbstract/sanskrit-sandhi-boundaries-v2.adaption-agri-qa-with-evidence-boundaries
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-agri_qa_with_evidence_boundaries
This dataset contains question-answer pairs focused on agricultural best practices, including crop management, pest control, and irrigation strategies. Each entry provides a direct, evidence-based response followed by a clearly defined 'evidence boundary' that limits the scope of the advice and advises verification with local conditions. The content… See the full description on the dataset page: https://huggingface.co/datasets/yeziR4/adaption-agri-qa-with-evidence-boundaries.alea-legal-benchmark-sentence-paragraph-boundaries
ALEA Legal Benchmark: Sentence and Paragraph Boundaries
Note: This dataset is derived from the ALEA Institute's KL3M Data Project. It builds upon the copyright-clean training resources while adding specific boundary annotations for sentence and paragraph detection.
Description
This dataset provides a comprehensive benchmark for sentence and paragraph boundary detection in legal documents. It was developed to address the unique challenges legal text poses for standard… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/alea-legal-benchmark-sentence-paragraph-boundaries.streamsafe-500-boundaries
StreamSafe 500 Boundary Annotations
This is a 500-trace streaming-boundary dataset derived from
StreamSafe. It contains
deterministic trace identifiers, prefix boundary positions, labels, harm categories, and
sentence-level unsafe-onset intervals. It includes the unchanged source prompt and full assistant response so it can be used without separately downloading StreamSafe. Prefix strings are reconstructed as response[:end_character].
Content warning
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vovarus12go/streamsafe-500-boundaries.
