CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01prem-research /spider Spider Unified dataset Documentation comming soon question-answering100M<n<1B1 likes1.4k downloads2y agoHugging Face02Spiderman01 /MedPriv-Bench_dataset MedPriv-Bench MedPriv-Bench evaluates the privacy--utility trade-off of language models in medical open-ended question answering. Each example contains a synthetic patient context, injected privacy-sensitive facts, a question, and a ground-truth answer. The benchmark supports evaluating whether a model remains clinically useful while avoiding disclosure of the injected facts. Split sizes Split Rows Composition train 2,200 1,315 benchmark-construction… See the full description on the dataset page: https://huggingface.co/datasets/Spiderman01/MedPriv-Bench_dataset.tabularquestion-answering1K<n<10K0 likes51 downloads1mo agoHugging Face03lianghsun /spider-text2sql-bench Dataset Card for spider-text2sql-bench spider-text2sql-bench 是 Spider 1.0 官方訓練集之 OpenAI Messages 格式版本,共 7,000 筆,將原始之 question / schema / sql 重新組裝為 system / user / assistant 三 role 之對話結構。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可作為 Text-to-SQL 模型之 SFT 訓練語料,亦可直接用於 benchmark evaluation pipeline(以 user 作為 prompt,比對模型輸出與 assistant 之標準答案 SQL)。 Dataset Details Dataset Description Spider 1.0 為 Yale LILY Group 於 EMNLP 2018 發表之大規模跨領域 Text-to-SQL… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/spider-text2sql-bench.texttext-generation1K<n<10K0 likes35 downloads5mo agoHugging Face04Porameht /spider_th Spider Thai Dataset Thai translation of the official Spider benchmark (A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task). Dataset Description This dataset contains Thai translations of the Spider text-to-SQL benchmark, translated from the official Spider data source. Source Original Dataset: Spider Benchmark Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/spider_th.texttable-question-answering1K<n<10K0 likes23 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.