datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hailuo-ai-voices
Hailuo AI Voices Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
📊 Dataset Overview
The dataset provides a comprehensive collection of voice samples with the following features:
Feature
Description
Audio Files
High-quality WAV format recordings
Transcription
Accurate transcriptions of each… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-voices.hailuo-ai-jokes
Hailuo AI Jokes Dataset 🎤
A curated collection of high-quality voice recordings with corresponding transcriptions and phoneme analysis. This dataset is designed for speech recognition, text-to-speech, and voice analysis tasks.
🎙️ Dataset Content
The dataset contains a diverse set of synthetic voice recordings generated by Hailuo AI Audio. The texts are sourced from a variety of public domain jokes and humorous anecdotes. Each audio sample is accompanied by the… See the full description on the dataset page: https://huggingface.co/datasets/unlimitedbytes/hailuo-ai-jokes.unlimitedfafnir
Bangumi Image Base of Unlimited Fafnir
This is the image base of bangumi Unlimited Fafnir, we detected 17 characters, 1386 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/unlimitedfafnir.google-cloud-voice-mixmultiprovider-voice-mixunlimited-torture12-rununlimited-ocr-multipage-smoke
Document OCR using Unlimited-OCR
This dataset contains OCR results for davanstrien/unlimited-ocr-pdf-test
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: davanstrien/unlimited-ocr-pdf-test
Model: baidu/Unlimited-OCR
Mode: multi-page (one document per row)
Number of Samples: 2
Processing Time: 2.3 min
Processing Date: 2026-06-28 09:20 UTC
Output Column: markdown
Max Pages/Document: 20
Split: train
Output
The column… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/unlimited-ocr-multipage-smoke.longalpaca_1k_unlimited_testDataset preprocessed from https://huggingface.co/datasets/Yukang/LongAlpaca-12k.
This contains 1000 samples that have a minimum length of 16k tokens.
Script to reproduce
from datasets import load_dataset
from transformers import AutoTokenizer
import pandas as pd
import pyarrow as pa
import pyarrow.parquet as pq
# Load the dataset and tokenizer
data = load_dataset("Yukang/LongAlpaca-12k")
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.1", trust_remote_code=True)… See the full description on the dataset page: https://huggingface.co/datasets/casperhansen/longalpaca_1k_unlimited_test.unlimited-ocr-britannica-smoke
Document OCR using Unlimited-OCR
This dataset contains OCR results for davanstrien/encyclopaedia-britannica-1771
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: davanstrien/encyclopaedia-britannica-1771
Model: baidu/Unlimited-OCR
Mode: single image per row
Number of Samples: 8
Processing Time: 3.0 min
Processing Date: 2026-06-28 09:17 UTC
Output Column: markdown
Split: train
Output
Grounding markup was stripped… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/unlimited-ocr-britannica-smoke.dogovors-unlimited-ocr
Dogovors Unlimited-OCR Dataset
OCR/document-layout dataset prepared for fine-tuning baidu/Unlimited-OCR.
Files
train.jsonl contains one JSON object per document.
images/ contains the page images referenced by relative path.
JSONL Schema
{
"images": [
"images/doc_001_page_001.jpg",
"images/doc_001_page_002.jpg"
],
"question": "Multi page parsing.",
"answer": "<PAGE><|det|>title [400, 60, 660, 73]<|/det|>...
<PAGE><|det|>text [100… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/dogovors-unlimited-ocr.Unlimited-Creativity-Chain-of-Thoughtunlimited-ocr-archive-testUncensored_and_unlimitedunlimited-ocr-pdf-testunlimited-ocr-smoke-v2
Document OCR using Unlimited-OCR
This dataset contains OCR results for davanstrien/unlimited-ocr-smoke-v2
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: davanstrien/unlimited-ocr-smoke-v2
Model: baidu/Unlimited-OCR
Number of Samples: 5
Processing Time: 2.2 min
Processing Date: 2026-06-30 13:49 UTC
Output Column: markdown
Split: train
Output
The column holds the model's raw layout-grounded markdown: text spans tagged… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/unlimited-ocr-smoke-v2.unlimited-ocr-smoke
Document OCR using Unlimited-OCR
This dataset contains OCR results for davanstrien/ufo-ColPali
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: baidu/Unlimited-OCR
Mode: single image per row
Number of Samples: 5
Processing Time: 2.3 min
Processing Date: 2026-06-28 09:12 UTC
Output Column: markdown
Split: train
Output
The column holds the model's raw layout-grounded markdown: text spans… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/unlimited-ocr-smoke.unlimited-ocr-grounding-test
Document OCR using Unlimited-OCR
This dataset contains OCR results for davanstrien/ufo-ColPali
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: davanstrien/ufo-ColPali
Model: baidu/Unlimited-OCR
Number of Samples: 3
Processing Time: 2.4 min
Processing Date: 2026-06-28 13:46 UTC
Output Column: markdown
Split: train
Output
Grounding markup was stripped (--strip-grounding); the column holds clean text.
Tables are returned… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/unlimited-ocr-grounding-test.dmb-bakeoff2-unlimitedunlimitedhucker-bakeoff2-unlimited-ordinary
Document OCR using Unlimited-OCR
This dataset contains OCR results for bokane/hucker-ocr-bakeoff
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: bokane/hucker-ocr-bakeoff
Model: baidu/Unlimited-OCR
Number of Samples: 12
Processing Time: 3.2 min
Processing Date: 2026-08-13 22:28 UTC
Output Column: markdown
Split: train
Output
Grounding markup was stripped (--strip-grounding); the column holds clean text.
Tables are… See the full description on the dataset page: https://huggingface.co/datasets/bokane/hucker-bakeoff2-unlimited-ordinary.hucker-bakeoff2-unlimited-hard
Document OCR using Unlimited-OCR
This dataset contains OCR results for bokane/hucker-ocr-bakeoff-hard
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: bokane/hucker-ocr-bakeoff-hard
Model: baidu/Unlimited-OCR
Number of Samples: 8
Processing Time: 2.7 min
Processing Date: 2026-08-13 22:31 UTC
Output Column: markdown
Split: train
Output
Grounding markup was stripped (--strip-grounding); the column holds clean text.… See the full description on the dataset page: https://huggingface.co/datasets/bokane/hucker-bakeoff2-unlimited-hard.unlimited-shake
Merge Fellas Mod APK: La Guía Definitiva del Puzzle Más Adictivo
➡️ ¡Descarga la última versión de Merge Fellas Mod APK ahora! Haz clic aquí:
https://modhello.com/es/merge-fellas/
Merge Fellas – El Juego de Puzzle Definitivo para Amantes de Mascotas y Memes
Merge Fellas Mod APK es uno de los juegos de puzzle más recientes y entretenidos disponibles para usuarios de Android. Con una jugabilidad relajante, gráficos encantadores y una mecánica de fusión creativa, se… See the full description on the dataset page: https://huggingface.co/datasets/merge-fellas/unlimited-shake.unlimited-ocr-windowssecond-model-our-captions
Dataset Card for "second-model-our-captions"
More Information needed
mqw-bakeoff2-unlimited-hard
Document OCR using Unlimited-OCR
This dataset contains OCR results for bokane/mqw-bakeoff-hard
produced by baidu/Unlimited-OCR with vLLM.
Processing Details
Source Dataset: bokane/mqw-bakeoff-hard
Model: baidu/Unlimited-OCR
Number of Samples: 6
Processing Time: 2.8 min
Processing Date: 2026-08-14 04:14 UTC
Output Column: markdown
Split: train
Output
Grounding markup was stripped (--strip-grounding); the column holds clean text.
Tables are returned… See the full description on the dataset page: https://huggingface.co/datasets/bokane/mqw-bakeoff2-unlimited-hard.
