datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic_relations_extraction
Dataset Card for "Semantic Relations Extraction"
Dataset Description
Repository
The "Semantic Relations Extraction" dataset is hosted on the Hugging Face platform, and was created with code from this GitHub repository.
Purpose
The "Semantic Relations Extraction" dataset was created for the purpose of fine-tuning smaller LLama2 (7B) models to speed up and reduce the costs of extracting semantic relations between entities in texts. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/DehydratedWater42/semantic_relations_extraction.concept-to-root-dictionary
🌿 Concept-to-Root Dictionary
A mapping of universal concepts to Arabic triliteral roots for semantic compression
📖 Overview
This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models.
What are Arabic Roots?
Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots:
Root
Core Meaning
Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.Syntactic-Semantic-Annotated-Italian-Corpus
Annotazione Sintattico-Funzionale e Disambiguazione della Lingua Italiana
This dataset was generated by fetching random first paragraphs from Italian Wikipedia (it.wikipedia.org)
and then processing them using Gemini AI with the following goal:
Processing Goal: riduci la ambiguità aggiungi tag grammaticali (soggetto) (verbo) eccetera. e tag funzionali es. (indica dove è nato il soggetto) (indica che il soggetto possiede l'oggetto) eccetera
Source Language: Italian (from Wikipedia)… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Syntactic-Semantic-Annotated-Italian-Corpus.mlsif-llm-semantic-integrity
MLSIF: Multi-Layer Semantic Integrity Framework Evaluation Dataset
Dataset Summary
This dataset accompanies the paper "Multi-Layer Semantic Integrity Framework for Knowledge-Reliable and Consistent Responses in Large Language Models" (Abishethvarman, Sabrina & Kwan). It contains the prompt set, raw model responses, and per-response evaluation scores used to benchmark eight open-source LLMs on semantic integrity, using a novel Semantic Integrity Index (SII).
The… See the full description on the dataset page: https://huggingface.co/datasets/abishethvarman/mlsif-llm-semantic-integrity.
