datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embedding-models
Reference models for integration into HF for Legal 🤗
This dataset comprises a collection of models aimed at streamlining and partially automating the embedding process. Each model entry within this dataset includes essential information such as model identifiers, embedding configurations, and specific parameters, ensuring that users can seamlessly integrate these models into their workflows with minimal setup and maximum efficiency.
Dataset Structure
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/HFforLegal/embedding-models.ft-embeddingmodel-RAG-dataset
Dataset Card for Dataset Name
This dataset aims to be a base template for fine-tuning embedding models for enhanced retrieval performance in RAG pipelines.
It has been generated locally using Mistral:7B on Ollama using a simple prompt that prompts the model to generate 5 questions for each document chunk of Apple's Environmental Progress Report 2024
Dataset Details
Dataset Description
Curated by: Likhit Juttada
Funded by [optional]: NA
Credits… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/ft-embeddingmodel-RAG-dataset.
