dh-unibe/image-text_zh-regierungsratsprotokolle
Dataset Card for image-text_zh-regierungsratsprotokolle This dataset was created using pagexml-hf converter from Transkribus PageXML data. Dataset Summary This dataset contains 152786 images. The document includes images and transcriptions of Regierungsratsprotokolle, transcribed within the frame of a project by the State Archives of Zurich. For more information see: (https://www.zentraleserien.zh.ch/home) Geographical scope: SwitzerlandPeriod: 1803-1883Languages:… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_zh-regierungsratsprotokolle.
Dataset Card for image-text_zh-regierungsratsprotokolle
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 152786 images.
The document includes images and transcriptions of Regierungsratsprotokolle, transcribed within the frame of a project by the State Archives of Zurich. For more information see: (https://www.zentraleserien.zh.ch/home)
Geographical scope: Switzerland<br>Period: 1803-1883<br>Languages: German<br>Type of document: Protocols<br>Provenance: State Archives of Zurich<br>
The Following Signatures are Included
- MM1001
- MM1002
- MM1003
- MM1004
- MM1005
- MM1006
- MM1007
- MM1008
- MM1009
- MM1010
- MM1011
- MM1012
- MM1013
- MM1014
- MM1015
- MM1016
- MM1017
- MM1018
- MM1019
- MM1020
- MM1021
- MM1022
- MM1023
- MM1024
- MM1025
- MM1026
- MM1027
- MM1028
- MM1029
- MM1030
- MM1031
- MM1032
- MM1033
- MM1034
- MM1035
- MM1036
- MM1037
- MM1038
- MM1039
- MM1040
- MM1041
- MM1042
- MM1043
- MM1044
- MM1045
- MM1046
- MM1047
- MM1048
- MM1049
- MM1050
- MM1051
- MM1052
- MM1053
- MM1054
- MM1055
- MM1056
- MM1057
- MM1058
- MM1059
- MM1060
- MM1061
- MM1062
- MM1063
- MM1064
- MM1065
- MM1066
- MM1067
- MM1068
- MM1069
- MM1070
- MM1071
- MM1072
- MM1073
- MM1074
- MM1075
- MM1076
- MM1077
- MM1078
- MM1079
- MM1080
- MM1081
- MM1082
- MM1083
- MM1084
- MM1085
- MM1086
- MM1087
- MM1088
- MM1089
- MM1090
- MM1091
- MM1092
- MM1093
- MM1094
- MM1095
- MM1096
- MM1097
- MM1098
- MM1099
- MM1100
- MM1101
- MM1102
- MM1103
- MM1104
- MM1105
- MM1106
- MM1107
- MM1108
- MM2101
- MM2102
- MM2103
- MM2104
- MM2105
- MM2106
- MM2107
- MM2108
- MM2109
- MM2110
- MM2111
- MM2112
- MM2113
- MM2114
- MM2115
- MM2116
- MM2117
- MM2118
- MM2119
- MM2120
- MM2121
- MM2122
- MM2123
- MM2124
- MM2126
- MM2127
- MM2128
- MM2129
- MM2130
- MM2131
- MM2132
- MM2133
- MM2134
- MM2135
- MM2136
- MM2137
- MM2139
- MM2140
- MM2141
- MM2142
- MM2143
- MM2144
- MM2145
- MM2146
- MM2147
- MM2148
- MM2149
- MM2150
- MM2151
- MM2152
- MM2153
- MM2154
- MM2155
- MM2156
- MM2157
- MM2158
- MM2159
- MM2160
- MM2161
- MM2162
- MM2163
- MM2164
- MM2165
- MM2166
- MM2167
- MM2168
- MM2169
- MM2170
- MM2171
- MM2172
- MM2173
- MM2174
- MM2175
- MM2176
- MM2177
- MM2178
- MM2179
- MM2180
- MM2181
- MM2182
- MM2183
- MM2184
- MM2185
- MM2186
- MM2187
- MM2188
- MM2189
- MM2190
- MM2191
- MM2192
- MM2193
- MM2194
- MM2195
- MM2196
- MM2197
- MM2198
- MM2199
- MM2200
- MM2201
- MM2202
- MM2203
- MM2204
- MM2205
- MM2206
- MM2207
- MM2208
- MM2209
- MM2210
- MM2211
- MM2212
- MM2213
- MM2214
- MM2215
- MM2216
- MM2217
- MM2218
- MM2219
- MM2220
- MM2221
- MM2222
- MM2223
- MM2224
- MM2225
- MM2226
- MM2227
- MM2228
- MM2229
- MM2230
- MM2231
- MM2232
- MM2233
- MM2234
- MM2235
- MM2236
- MM2237
- MM2238
- MM2239
- MM2240
- MM2241
- MM2242
- MM2243
- MM2244
- MM2247
- MM2248
- MM2249
- MM2251
- MM2253
- MM2254
Dataset Structure
Data Splits
- train: 152786 samples
Dataset Size
- Approximate total size: 1783572.44 MB
- Total samples: 152786
Features
- image:
Image(mode=None, decode=False) - xml_content:
Value('string') - filename:
Value('string') - project_name:
Value('string')
Data Organization
Data is organized as parquet shards by split and project:
data/
├── <split>/
│ └── <project_name>/
│ └── <timestamp>-<shard>.parquetThe HuggingFace Hub automatically merges all parquet files when loading the dataset.
Usage
from datasets import load_dataset
# Load entire dataset
dataset = load_dataset("dh-unibe/transkribus-exports-11481-raw-xml")
# Load specific split
train_dataset = load_dataset("dh-unibe/transkribus-exports-11481-raw-xml", split="train")