datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parsebench-recipe-runs
ParseBench recipe runs
Experiment log for document-parsing runs on the public ParseBench test
subset (llamaindex/ParseBench), scored with the official open-source
ParseBench evaluator (run-llama/ParseBench).
Recipes combine CLI coding agents doing vision parsing with deterministic
PDF text-layer tools (word-bbox snapping, style extraction from span flags
and vector-drawing geometry).
Sample: official test subset - 12 single-page PDFs, 3 per category
(chart / layout / table /… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/parsebench-recipe-runs.bol-bench
bol-bench: a bill-of-lading parsing benchmark
25 synthetic single-page bill-of-lading PDFs with exact field-level ground
truth, built to benchmark PDF parsers on logistics documents. To our knowledge
this is the first public bill-of-lading parsing benchmark — Hugging Face
previously had no BoL parser or dataset, and ParseBench
(arXiv:2604.08538) contains no logistics
documents.
Leaderboard + methodology:
okrapdf.com/blog/bill-of-lading-ocr-benchmark
— 24 parsers scored… See the full description on the dataset page: https://huggingface.co/datasets/sleepyheeler/bol-bench.Wattpad-metadata-hotqwen3.6-27b-self-data-distillation-dataset
Qwen3.6-27B Self-Data-Distillation Trajectories
Single‑turn reasoning trajectories generated by running Qwen3.6‑27B (via vLLM). Each trajectory contains a system prompt, a user task, and the model's full output (including reasoning steps embedded in the assistant content field).
Data Format
Four JSONL files, one per category. Each line is:
{
"id": "traj_<timestamp>_<idx>_<seq>",
"source": "synthetic-qwen3.6-27b",
"task": "<the prompt given to the model>"… See the full description on the dataset page: https://huggingface.co/datasets/sleepyeldrazi/qwen3.6-27b-self-data-distillation-dataset.Wattpad-metadata-newRdiffusionRdiffusion
We're releasing the entire corpus of publicly available songs from the Riffusion platform—generated and shared by their user community. Through extensive scraping of their exposed API, we’ve collected over 2,000 artificial songs, including every metadata field and downloadable asset that was accessible at the time.
📦 Included Data
Audio & Visual Assets:audio_url, audio_b64, image_url, image_b64, video_url
Metadata & Structure:id, title, author_id, created_at… See the full description on the dataset page: https://huggingface.co/datasets/sleeping-ai/Rdiffusion.PlayAI-VoiceExcited to share Play AI Voice Profile. We release 267 unique voice profiles including Israeli, Arabic, Russian, Filipino and many other exclusive voice profiles. Play AI was recently acquired by Meta which sparked our interest in releasing this dataset.
