hanse
Datasets
All datasets matching “hanse”HanselHansel is a high-quality human-annotated Chinese entity linking (EL) dataset, used for testing Chinese EL systems' generalization ability to tail entities and emerging entities.
The test set contains Few-shot (FS) and zero-shot (ZS) slices, has 10K examples and uses Wikidata as the corresponding knowledge base.
The training and validation sets are from Wikipedia hyperlinks, useful for large-scale pretraining of Chinese EL systems.LRID
LRID Dataset
📖 Overview
This is the full version of the Low-light Raw Image Denoising (LRID) dataset, designed for low-light image denoising research. It is part of the PMN_TPAMI project (Learnability Enhancement for Low-light Raw Denoising: A Data Perspective, TPAMI 2024).
GitHub: https://github.com/megvii-research/PMN/tree/TPAMI
🗂️ Dataset Structure
PMN_TPAMI
└─LRID # Full LRID Dataset (Raw Data)
├─bias # Dark Frames
├─bias-hot #… See the full description on the dataset page: https://huggingface.co/datasets/hansen97/LRID.hansen-catalogoMMedFD
MMedFD: A Real-World Healthcare Benchmark for Multi-Turn Full-Duplex Automatic Speech Recognition
⚠️Data Availability
Full access requires internal approval and a research-only data use agreement.
🚫 Non-Commercial Use
This dataset is provided for non-commercial research and education only. Commercial use is prohibited.Researchers who wish to request full access may contact us with a brief description of their affiliation, project goals, intended use, and data… See the full description on the dataset page: https://huggingface.co/datasets/HanselZz/MMedFD.medmcqa-filtered-v16
MedMCQA Filtered v16
This commit adds normalized source explanations to official validation for a
mixed gold-rationale-assisted likelihood evaluation. The immutable training
revision remains 29cf55f62c523dc9c1c45e2d77f52208058e7d19; evaluation clients must pin the
full SHA of this commit separately.
Splits
Split
Rows
Contract
train
63,316
Byte-logically unchanged from the training revision
tuning
5,000
Byte-logically unchanged from the training… See the full description on the dataset page: https://huggingface.co/datasets/hanseungwook/medmcqa-filtered-v16.hanse-kurrent-xvi-rawxml
Dataset Card for hanse-kurrent-xvi-rawxml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
The dataset contains transkriptions of the Minutes of the low German town assemblies (Niederdeutsche Städtetage) from the 16th Century from various archives. The Texts include middle low German and New High German Languages
Transcriptions were created according to the following guidelines: Forschungsstelle für die Geschichte der Hanse und des Ostseeraums… See the full description on the dataset page: https://huggingface.co/datasets/fgho/hanse-kurrent-xvi-rawxml.
