datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Test_Semantic_Searchlab05-semantic-searchsemantic_searchsemantic-job-search-dataset
Semantic Job Search Dataset (Synthetic)
Overview
This dataset contains 10,000 synthetic job postings designed for semantic search.Each record represents a job listing with structured attributes (e.g., field, location, job type) plus a short natural-language description.
The dataset was generated as part of a course final project and is used to support a Gradio app that performs semantic job search using Sentence-Transformer embeddings and cosine similarity.… See the full description on the dataset page: https://huggingface.co/datasets/aurele1/semantic-job-search-dataset.lab05-semantic-searchlab05-semantic-searchlab05-semantic-search
Speak the Patient's Language: semantic search vs keyword search
Completed by James Seegel for MIS 752 at UNLV.
The teaching corpus and evaluation queries were provided in Dr. Richard Young's lab notebook.
My takeaways
Sentences within the same topic averaged 0.235 cosine similarity, compared with 0.129 across topics, showing some separation between subjects. Keyword search handled “my refill is not ready at the pharmacy” because “refill,” “ready,” and “pharmacy”… See the full description on the dataset page: https://huggingface.co/datasets/JamesSeegel/lab05-semantic-search.lab05-semantic-searchlab05-semantic-searchsemantic_search_t5_format2
en en en doğru dataset
kw'ler full
semantic_search_t5_format3semantic-search-channelssemantic_search_t5_format5semantic_search_t5_formatsemantic-search-data
