datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
token-counting-edge-cases
token-counting-edge-cases
20 short strings with approximate token counts across three tokenizer families: Claude, GPT (cl100k_base), and Llama (SentencePiece). Built for sanity-checking token counters / chunkers / context-window fitters.
The numbers are approximate — exact counts depend on tokenizer version, BOS/EOS handling, and surrounding context. Expect ±1–2 token jitter. Use these to catch order-of-magnitude bugs (e.g. "your counter says 200 tokens for one emoji"), not as… See the full description on the dataset page: https://huggingface.co/datasets/mukunda1729/token-counting-edge-cases.screenplay-format-edge-cases
Screenplay Format Edge Cases
48 original Fountain specimens in 24 contrast pairs — two near-identical inputs
per pair, at the points where the Fountain syntax leaves a choice. In 17 pairs
the one difference changes how the lines are classified. In the other 7 it
changes the surface and the labels hold: a lowercase scene prefix, a cue
extension, a non-Latin cue, escaped characters, a dual-dialogue caret, an inline
note and centered-text markers.
Version: 1.0.0 · Maintainer:… See the full description on the dataset page: https://huggingface.co/datasets/Inkwell-Software/screenplay-format-edge-cases.
