SciKnowOrg/ontolearner-materials_science_and_engineering
Materials Science And Engineering Domain Ontologies Overview Materials Science and Engineering is a multidisciplinary domain that focuses on the study and application of materials, emphasizing their structure, properties, processing, and performance in engineering contexts. This field is pivotal for advancing knowledge representation, as it integrates principles from physics, chemistry, and engineering to innovate and optimize materials for diverse technological… See the full description on the dataset page: https://huggingface.co/datasets/SciKnowOrg/ontolearner-materials_science_and_engineering.
license: mit language:
- en tags:
- OntoLearner
- Ontology-Learning
- Materials-Science-And-Engineering
- Benchmark-Ontologies pretty_name: Materials Science And Engineering ---
<div align="center"> <img src="https://raw.githubusercontent.com/sciknoworg/OntoLearner/main/images/logo.png" alt="OntoLearner" style="display: block; margin: 0 auto; width: 500px; height: auto;"> <h1 style="text-align: center; margin-top: 1em;">Materials Science And Engineering Domain Ontologies</h1> <a href="https://github.com/sciknoworg/OntoLearner"><img src="https://img.shields.io/badge/GitHub-OntoLearner-blue?logo=github" /></a> </div>
Overview
Materials Science and Engineering is a multidisciplinary domain that focuses on the study and application of materials, emphasizing their structure, properties, processing, and performance in engineering contexts. This field is pivotal for advancing knowledge representation, as it integrates principles from physics, chemistry, and engineering to innovate and optimize materials for diverse technological applications. By systematically categorizing and modeling material-related data, this domain facilitates the development of new materials and enhances the understanding of their behavior under various conditions.
Dataset Files
Each ontology directory contains the following files:
<ontology_id>.<format>- The original ontology fileterm_typings.json- Dataset of term to type mappingstaxonomies.json- Dataset of taxonomic relationsnon_taxonomic_relations.json- Dataset of non-taxonomic relations<ontology_id>.rst- Documentation describing the ontology
Usage
These datasets are intended for ontology learning research and applications. Here's how to use them with OntoLearner:
First of all, install the OntoLearner library via PiP:
pip install ontolearnerHow to load an ontology or LLM4OL Paradigm tasks datasets?
from ontolearner import BattINFO
ontology = BattINFO()
# Load an ontology.
ontology.load()
# Load (or extract) LLMs4OL Paradigm tasks datasets
data = ontology.extract()How use the loaded dataset for LLM4OL Paradigm task settings?
# Import core modules from the OntoLearner library
from ontolearner import BattINFO, LearnerPipeline, train_test_split
# Load the BattINFO ontology, which contains concepts related to wines, their properties, and categories
ontology = BattINFO()
ontology.load() # Load entities, types, and structured term annotations from the ontology
ontological_data = ontology.extract()
# Split instances into train and test sets
train_data, test_data = train_test_split(ontological_data, test_size=0.2, random_state=42)
# Initialize a multi-component learning pipeline (retriever + LLM)
# This configuration enables a Retrieval-Augmented Generation (RAG) setup
pipeline = LearnerPipeline(
retriever_id='sentence-transformers/all-MiniLM-L6-v2', # Dense retriever model for nearest neighbor search
llm_id='Qwen/Qwen2.5-0.5B-Instruct', # Lightweight instruction-tuned LLM for reasoning
hf_token='...', # Hugging Face token for accessing gated models
batch_size=32, # Batch size for training/prediction if supported
top_k=5 # Number of top retrievals to include in RAG prompting
)
# Run the pipeline: training, prediction, and evaluation in one call
outputs = pipeline(
train_data=train_data,
test_data=test_data,
evaluate=True, # Compute metrics like precision, recall, and F1
task='term-typing' # Specifies the task
# Other options: "taxonomy-discovery" or "non-taxonomy-discovery"
)
# Print final evaluation metrics
print("Metrics:", outputs['metrics'])
# Print the total time taken for the full pipeline execution
print("Elapsed time:", outputs['elapsed_time'])
# Print all outputs (including predictions)
print(outputs)For more detailed documentation, see the 
Citation
If you find our work helpful, feel free to give us a cite.
@inproceedings{babaei2023llms4ol,
title={LLMs4OL: Large language models for ontology learning},
author={Babaei Giglou, Hamed and D’Souza, Jennifer and Auer, S{"o}ren},
booktitle={International Semantic Web Conference},
pages={408--427},
year={2023},
organization={Springer}
}