datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hiro-Pharma-RAG-Benchmark
Hiro Pharma RAG Benchmark
This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine.
The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references.
Repository Contents
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.sft_alfworld_v5_nocot_nomarker
ALFWorld v5 Trajectory Dataset (No CoT, No Marker)
An ALFWorld household-task agent trajectory dataset in OpenAI chat messages format, prepared for supervised fine-tuning (SFT).Derived from u-10bei/sft_alfworld_trajectory_dataset_v5 by removing Chain-of-Thought lines and stripping the Act: marker from assistant outputs.
Splits
train: 2,376
eval: 126
Stratified split (95/5) by metadata.task_type and metadata.difficulty with seed=3407 (rare strata <10 grouped as… See the full description on the dataset page: https://huggingface.co/datasets/hiro0904/sft_alfworld_v5_nocot_nomarker.nick_name_from_hiroiki-ariyoshi
data format
{
"name": "person_name",
"nickname": "person_nickname"
}
