datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
p4_datasetp4datasetdogovors-unlimited-ocr
Dogovors Unlimited-OCR Dataset
OCR/document-layout dataset prepared for fine-tuning baidu/Unlimited-OCR.
Files
train.jsonl contains one JSON object per document.
images/ contains the page images referenced by relative path.
JSONL Schema
{
"images": [
"images/doc_001_page_001.jpg",
"images/doc_001_page_002.jpg"
],
"question": "Multi page parsing.",
"answer": "<PAGE><|det|>title [400, 60, 660, 73]<|/det|>...
<PAGE><|det|>text [100… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/dogovors-unlimited-ocr.mcp-circuit-datasetqwen_ds_russian_v1
Qwen-DS-Russian v1
A Russian-language dataset focused on Data Science, Machine Learning, and Python data manipulation.
Dataset Structure
Format: Instruction-Response pairs.
Samples: 2,044 curated Q&A pairs.
Topics:
General Python for Data Science
Pandas & NumPy operations
Classic Machine Learning (Scikit-Learn)
Data Visualization (Matplotlib, Seaborn)
Deep Learning foundations
Use Case
Ideal for fine-tuning small-to-medium LLMs to provide… See the full description on the dataset page: https://huggingface.co/datasets/p4ulbr4dl3y/qwen_ds_russian_v1.qwen_ds_v4_gold_datasetp4-datasetP_4_survey_16_3Bp4datasetsub_175_p_p4storyPro_p4storyPro_p40storyPro_p41storyPro_p42storyPro_p43storyPro_p44storyPro_p45storyPro_p46storyPro_p47storyPro_p49ab_552_p4ab_552_p40ab_552_p41ab_552_p42ab_552_p43ab_552_p44ab_552_p45ab_552_p46ab_552_p47ab_552_p48
