datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task960_ancora-ca-ner_named_entity_recognition
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task960_ancora-ca-ner_named_entity_recognition
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task960_ancora-ca-ner_named_entity_recognition.Home-Assistant-requests-for-intent-detection-and-function-recognition
Home Assistant Requests V2 Dataset
This dataset contains a list of requests and responses for a user interacting with a personal assistant that controls an instance of Home Assistant.
The updated V2 of the dataset is now multilingual, containing data in English, German, French, Spanish, and Polish. The dataset also contains multiple "personalities" for the assistant to respond in, such as a formal assistant, a sarcastic assistant, and a friendly assistant. Lastly, the dataset has… See the full description on the dataset page: https://huggingface.co/datasets/DaftP/Home-Assistant-requests-for-intent-detection-and-function-recognition.Pattern-Recognition
Pattern Completion Dataset
A 30 GB synthetic dataset of numeric sequence‑completion prompts and their next values, designed to teach large language models how to recognize and extrapolate patterns.
Each row contains a prompt (the sequence with a ? indicating the missing next element) and a completion (the correct next number).
Dataset Structure
Format: CSV (no header row)
Columns:
prompt – "Find the next number in the sequence: a,b,c,... ,?"
completion – the… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Pattern-Recognition.Pattern-Recognition
Pattern Completion Dataset
A 30 GB synthetic dataset of numeric sequence‑completion prompts and their next values, designed to teach large language models how to recognize and extrapolate patterns.
Each row contains a prompt (the sequence with a ? indicating the missing next element) and a completion (the correct next number).
Dataset Structure
Format: CSV (no header row)
Columns:
prompt – "Find the next number in the sequence: a,b,c,... ,?"
completion – the… See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Pattern-Recognition.task1544_conll2002_named_entity_recognition_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1544_conll2002_named_entity_recognition_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1544_conll2002_named_entity_recognition_answer_generation.law_entity_recognition
Dataset Card for Dataset Name
The dataset transforms complex legal passages into structured outputs, detailing entities, their interrelationships, and claims, providing a foundation for a legal knowledge graph to facilitate advanced analysis and applications.
Dataset Details
Dataset Description
The dataset in question is a specialized collection designed for legal text analysis, where each input is a passage of legal text—ranging from case law to statutory… See the full description on the dataset page: https://huggingface.co/datasets/rubenamtz0/law_entity_recognition.
