datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tr-legal-triplets
Turkish Legal QA Triplets (tr-legal-triplets)
A large-scale dataset of query/positive/negative triplets generated from Turkish
legal documents, designed for training and evaluating multilingual text embedding models on
Turkish legal content.
Subset
Records
Description
default
1,148,041
Full generated dataset
cleaned
1,043,913
Quality-filtered — recommended for training
Dataset Description
Source Data
The dataset is built from 19… See the full description on the dataset page: https://huggingface.co/datasets/yunus-emre/tr-legal-triplets.amadesus-trl-assistant-dataset-v2-0
AMADEUS_TRL_DATASET
Dataset Description
amadesu_trl_assistant_dataset is designed to train intelligent assistants in evaluating the Technology Readiness Level (TRL) in the field of agriculture, using the TRL metric developed by NASA. The dataset is organized into two parts:
Conceptual Knowledge Dataset: Provides essential knowledge about TRL concepts and definitions, levels, objectives, and goals for each level, as well as related technological development activities.… See the full description on the dataset page: https://huggingface.co/datasets/JsBetancourt/amadesus-trl-assistant-dataset-v2-0.kaggleds-corpus-task-based-search-bench
KaggleDS: Corpus for Task-based Dataset Search
KaggleDS is a benchmark corpus for evaluating task-driven dataset search — retrieving relevant tables from natural-language descriptions of analytical goals (e.g., "Analyze trends in the California real estate market over the past decade") rather than keyword or schema-level queries.
The corpus was introduced in the paper "DataForager: Enabling Flexible Need-Aligned Dataset Navigation".
Corpus Overview
Train… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/kaggleds-corpus-task-based-search-bench.
