datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
docling-nlp-datasetsThis repository contains the models used for docling-nlp.
Contents
This model repository packages the pretrained assets used by Docling’s NLP
components:
CRF models for material classification and English part-of-speech tagging
fastText models for language detection, metadata, semantic, topic, and person-name classification
Regular-expression assets for geographic-location extraction and unit handling
A default tokenizer model
Correct workflow to add new files… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/docling-nlp-datasets.test-eval-docs-docling-plain
PDF Document Processing with Docling
This dataset contains structured markdown extraction from PDFs in baobabtech/test-eval-documents
using Docling with hierarchical parsing.
Processing Details
Source Dataset: baobabtech/test-eval-documents
Number of PDFs: 20
Processing Time: 8.4 minutes
Processing Date: 2025-12-02 15:40 UTC
Configuration
PDF Column: pdf_bytes
Dataset Split: train
Dataset Structure
The dataset contains all original columns plus:… See the full description on the dataset page: https://huggingface.co/datasets/baobabtech/test-eval-docs-docling-plain.
