entity-dataset
dutch-law-enforcement-entity-resolution-dataset
Dutch Law Enforcement Entity Resolution Benchmark
A collection of synthetic benchmark datasets for entity resolution and record linkage research, modelled on real Dutch and EU/EEA law enforcement data schemas.
All data is 100% synthetic.
No records correspond to real persons, vehicles, phone numbers, transactions, or criminal histories. Every name, date, plate number, IBAN, IMSI, and identifier is procedurally generated. The datasets are freely usable, shareable, and… See the full description on the dataset page: https://huggingface.co/datasets/zal-analytics-core/dutch-law-enforcement-entity-resolution-dataset.GLAMI-Entity-Matching-Dataset
GLAMI Duplication Detection
Product-duplicate detection over GLAMI e-commerce listings: ~1.3M product
images plus multilingual titles, descriptions and attributes, with labelled
groups of items that do or do not refer to the same physical product.
Released under the Apache License 2.0 — see LICENSE.
TODO: describe how the labels were produced.
Structure
Config
Files
Contents
images
images/shard-*.parquet
itemId → image bytes, one row per product image… See the full description on the dataset page: https://huggingface.co/datasets/zidcenek/GLAMI-Entity-Matching-Dataset.task1448_disease_entity_extraction_ncbi_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1448_disease_entity_extraction_ncbi_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1448_disease_entity_extraction_ncbi_dataset.funding-entity-extraction-dataset-mix
Funding Entity Extraction Dataset Mix
Training and evaluation corpus for funding-entity extraction from full text. This dataset is the data mix used to fine-tune cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward on top of meta-llama/Llama-3.1-8B-Instruct in two stages: SFT on the mix described below, followed by GRPO with a hierarchical F0.5 reward.
Loading the dataset
from datasets import load_dataset
# Default config (degraded… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/funding-entity-extraction-dataset-mix.kaggle-entity-annotated-corpus-ner-datasetDate: 2022-07-10
Files: ner_dataset.csv
Source: Kaggle entity annotated corpus
notes: The dataset only contains the tokens and ner tag labels. Labels are uppercase.
About Dataset
from Kaggle Datasets
Context
Annotated Corpus for Named Entity Recognition using GMB(Groningen Meaning Bank) corpus for entity classification with enhanced and popular features by Natural Language Processing applied to the data set.
Tip: Use Pandas Dataframe to load dataset if using Python for… See the full description on the dataset page: https://huggingface.co/datasets/rjac/kaggle-entity-annotated-corpus-ner-dataset.task1449_disease_entity_extraction_bc5cdr_dataset
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1449_disease_entity_extraction_bc5cdr_dataset
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1449_disease_entity_extraction_bc5cdr_dataset.
