datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
royal_society_corpus_LLM_generated_metadataaugmented_dataset_llm_generated_NER
📚 Augmented LLM-Generated NER Dataset for Scholarly Text
🧠 Dataset Summary
This dataset contains synthetically generated academic text tailored for Named Entity Recognition (NER) in the software engineering domain. The synthetic data augments scholarly writing using large language models (LLMs), with entity consistency maintained via token preservation.
The dataset is generated by merging and rephrasing pairs of annotated sentences from scholarly papers using… See the full description on the dataset page: https://huggingface.co/datasets/psresearch/augmented_dataset_llm_generated_NER.test_generated_seqtrain_generated_seqgenerated_seqsgenerated_train_diff_v2LLMgenerated_fictive_medical_report_and_summaries_with_omissions_label_Fr_Healthcare
🏥 French Synthetic Medical Reports and Summaries with Omission Labels (Fr-Healthcare)
This dataset contains fictitious French medical reports, each paired with a summary and a binary label indicating whether the summary omits relevant factual content. It is designed solely for evaluating factual consistency and omission detection in Natural Language Processing, particularly in the medical domain. We must emphasize that all names, identifiers, dates, medical information, and any… See the full description on the dataset page: https://huggingface.co/datasets/AchOk78/LLMgenerated_fictive_medical_report_and_summaries_with_omissions_label_Fr_Healthcare.generated_train_diff
