CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes59 downloads7d agoHugging Face02Jainam-11 /symptom-based-disease-prediction-v2 Symptom-Based Disease Prediction Dataset (Confidence-Aware) – v2 📘 Overview Version 2 of the Symptom-Based Disease Prediction Dataset is an improved and more structured medical reasoning dataset designed for Large Language Models (LLMs) and AI-driven healthcare research. This version introduces: cleaner formatting, improved disease grouping, better confidence separation, enhanced consistency, and more fine-tuning-friendly outputs. Each sample presents symptoms in… See the full description on the dataset page: https://huggingface.co/datasets/Jainam-11/symptom-based-disease-prediction-v2.textquestion-answering100K<n<1M1 likes33 downloads4mo agoHugging Face03gviviano /frames-benchmark-predictions BENCH-04: N-8 Research Google FRAMES Benchmark Evaluation Team Designation: N-8 ResearchLead Author: Greg VivianoOrganization: N-8 ResearchEvaluated Dataset: google/frames-benchmark (824 Questions) Executive Summary This repository contains the prediction dataset generated by N-8 Research's Deterministic Context Architecture across all 824 multi-step enterprise reasoning questions in Google's official google/frames-benchmark. Performance Scorecard… See the full description on the dataset page: https://huggingface.co/datasets/gviviano/frames-benchmark-predictions.tabularquestion-answeringn<1K0 likes29 downloads2mo agoHugging Face04dnaihao /table-sft-eval-predictions 💾 Raw Predictions for "What Really Matters for Table LLMs?" This dataset contains the raw model outputs from the experiments in: Naihao Deng, Sheng Zhang, Henghui Zhu, Shuaichen Chang, Jiani Zhang, Alexander Hanbo Li, Chung-Wei Hang, Hideo Kobayashi, Yiqun Hu, Patrick Ng. What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects. Findings of EACL 2026. https://aclanthology.org/2026.findings-eacl.195/ 🗂️ Layout… See the full description on the dataset page: https://huggingface.co/datasets/dnaihao/table-sft-eval-predictions.texttext-generation100K<n<1M0 likes23 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.