datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brandvoice-marketing-briefs
BrandVoice Marketing Briefs
Generic AI writes like a robot. This trains it to write like a brand.
The loop is closed. The LoRA adapter
trained on this data scored +24.3% copy quality (7.0 to 8.7) with a 56% win rate against its
base, on Adaption's held-out judge. Dataset, weights, evaluation, and reproduction are all public.
A corpus of real marketing copy. 6,339 lines scraped from the live pages of 83 brands (the actual
Stripe, Liquid Death, Ramp, Duolingo copy)… See the full description on the dataset page: https://huggingface.co/datasets/manifesta/brandvoice-marketing-briefs.autoresearch-manim
Autoresearch Manim
Curated Manim code-generation examples exported from the autoresearch_manim_finetune pipeline.
Preview Gallery
Preview
Preview
Preview
Machine learning: attention plus residual mixing
Physics: boundary layer flow near a surface
Biology: neuron structure and signal direction
Finance: compound growth over time
Economics: production frontier tradeoff
Neuroscience: action potential phases
Summary
Focus:… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/autoresearch-manim.mande-ancient-treasures-de-grunne-van-dyke-2016
mande-ancient-treasures-de-grunne-van-dyke-2016
Dataset created with PDF2Dataset -- OCR + structure-aware chunking pipeline.
Dataset Summary
Metric
Value
Total chunks
268
Avg chars/chunk
722
Avg images/chunk
0.14
Source files
1
Duplicates removed
0
Quality filtered
6
Schema
Column
Type
Description
chunk_id
string
Unique identifier: filename_chunk_N
text
string
Raw markdown chunk with image refs
text_clean… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/mande-ancient-treasures-de-grunne-van-dyke-2016.
