CoolFace
Datasetpublic

dh-unibe/image-text_medieval-scripts_xiv-xv-xvi

Dataset Card for image-text_medieval-scripts_xiv-xv-xvi This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 548322 samples across 1 split(s). Geographical scope: BelgiumPeriod: 1350-1550Languages: FlemishType of document: ProtocolProvenance: State Archives in Leuven Projects Included Itinera Nova Parts of Charters from Königsfelden SAL7304_full SAL7305_full SAL7306_full SAL7307… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_medieval-scripts_xiv-xv-xvi.

sourceHugging Facemitupdated 5mo agoView on Hugging Face
1likes2.3kdownloads
Dataset Card

Dataset Card for image-textmedieval-scriptsxiv-xv-xvi

This dataset was created using pagexml-hf converter from Transkribus PageXML data.

Dataset Summary

This dataset contains 548322 samples across 1 split(s).

Geographical scope: Belgium<br>Period: 1350-1550<br>Languages: Flemish<br>Type of document: Protocol<br>Provenance: State Archives in Leuven<br>

Projects Included

  • —Itinera Nova
  • —Parts of Charters from Königsfelden
  • —SAL7304_full
  • —SAL7305_full
  • —SAL7306_full
  • —SAL7307
  • —SAL7307_full
  • —SAL7308
  • —SAL7309
  • —SAL7310
  • —SAL7310_full
  • —SAL7311
  • —SAL7311_full
  • —SAL7312-7313
  • —SAL7314
  • —SAL7315
  • —SAL7316
  • —SAL7316_full
  • —SAL7317
  • —SAL7317_full
  • —SAL7318
  • —SAL7318_full
  • —SAL7319
  • —SAL7320
  • —SAL7321
  • —SAL7322
  • —SAL7323
  • —SAL7323_full
  • —SAL7324
  • —SAL7324_full
  • —SAL7325
  • —SAL7325_full
  • —SAL7326
  • —SAL7326_full
  • —SAL7327
  • —SAL7328
  • —SAL7329
  • —SAL7330
  • —SAL7331
  • —SAL7332
  • —SAL7333
  • —SAL7334
  • —SAL7335
  • —SAL7336
  • —SAL7337
  • —SAL7338
  • —SAL7339
  • —SAL7340
  • —SAL7341
  • —SAL7342
  • —SAL7343
  • —SAL7344
  • —SAL7345
  • —SAL7346
  • —SAL7347
  • —SAL7349
  • —SAL7350
  • —SAL7351
  • —SAL7352
  • —SAL7353
  • —SAL7354
  • —SAL7355
  • —SAL7356
  • —SAL7357
  • —SAL7358
  • —SAL7359
  • —SAL7362
  • —SAL7363
  • —SAL7370
  • —SAL7375_full
  • —SAL7376_full
  • —SAL7382
  • —SAL7384
  • —SAL7385
  • —SAL7386
  • —SAL7387
  • —SAL7392
  • —SAL7698_full
  • —SAL7699_full
  • —SAL7700_full
  • —SAL7701_full
  • —SAL7702_full
  • —SAL7703_full
  • —SAL7704
  • —SAL7705
  • —SAL7706
  • —SAL7709
  • —SAL7714
  • —SAL7715
  • —SAL7716
  • —SAL7717
  • —SAL7718
  • —SAL7719
  • —SAL7720
  • —SAL7721
  • —SAL7722
  • —SAL7723
  • —SAL7724
  • —SAL7725
  • —SAL7726
  • —SAL7727
  • —SAL7728
  • —SAL7729
  • —SAL7730
  • —SAL7731
  • —SAL7731_full
  • —SAL7732
  • —SAL7734_full
  • —SAL7735_full
  • —SAL7736_full
  • —SAL7738_full
  • —SAL7739_full
  • —SAL7742_full
  • —SAL7743_full
  • —SAL7744_full
  • —SAL7745_full
  • —SAL7746_full
  • —SAL7747_full
  • —SAL7748_full
  • —SAL7750_full
  • —SAL7751_full
  • —SAL7762_full
  • —SAL7767_full
  • —SAL7982_full
  • —SAL7983_full
  • —SAL7984_full
  • —SAL7985_full
  • —SAL7986_full
  • —SAL7987_full
  • —SAL7988_full
  • —SAL7989_full
  • —SAL7990_full
  • —SAL7991_full
  • —SAL7992_full
  • —SAL7993_full
  • —SAL7994_full
  • —SAL8021_full
  • —SAL8022_full
  • —SAL8023_full
  • —SAL8024_full
  • —SAL8114_full
  • —SAL8115_full
  • —SAL8120_full
  • —SAL8121_full
  • —SAL8123_full
  • —SAL8131_full
  • —SAL8132_full
  • —SAL8135_full
  • —SAL8154_full
  • —SAL8336_full
  • —SAL8337_full
  • —SAL8338_full
  • —SAL8339_full
  • —SAL8340_full
  • —SAL8341_full
  • —SAL8342_full
  • —SAL8343_full
  • —SAL8344_full
  • —SAL8345_full
  • —SAL8347_full
  • —SAL8348_full
  • —SAL8349_full
  • —SAL8350_full
  • —TRAININGVALIDATIONSETHGBFTM450
  • —TRAININGVALIDATIONSETHGBFTM475
  • —TRAININGVALIDATIONSETHGBFTM495
  • —Thuner Missiven
  • —u-17_0059
  • —u-17_0060
  • —u-17006101
  • —u-17006102
  • —u-17_0065
  • —u-17_0075
  • —u-17_0083
  • —u-17_0103
  • —u-17_0104
  • —u-17_0126
  • —u-17_0149
  • —u-17_0151
  • —u-17_0152
  • —u-17_0159
  • —u-17_0179
  • —u-17_0185a
  • —u-17_0187a

Dataset Structure

Data Splits

  • —train: 548322 samples

Dataset Size

  • —Approximate total size: 6635119.97 MB
  • —Total samples: 548322

Features

  • —image: Image(mode=None, decode=False)
  • —xml_content: Value('string')
  • —filename: Value('string')
  • —project_name: Value('string')

Data Organization

Data is organized as parquet shards by split and project:

data/
├── <split>/
│   └── <project_name>/
│       └── <timestamp>-<shard>.parquet

The HuggingFace Hub automatically merges all parquet files when loading the dataset.

Usage

python
from datasets import load_dataset

# Load entire dataset
dataset = load_dataset("dh-unibe/image-text_medieval-scripts_xiv-xv-xvi") 

# Load specific split
train_dataset = load_dataset("dh-unibe/image-text_medieval-scripts_xiv-xv-xvi", split="train")