CoolFace
6 results

enterprise-benchmark

hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes116 downloads1d agoHugging FaceAbdulrahmankalil /enterprise-llm-inference-benchmarks-2026 🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide) A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments. 🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation) Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.tabulartext-generationn<1K1 likes88 downloads7d agoHugging Facepriteshloke /enterprise-operations-benchmark 🌐 Kepler Enterprise Operations Exception Benchmark (10k Dataset) Author: Kepler Operations Intelligence (https://www.getkeplerops.com) This is the canonical evaluation benchmark for AI models, LLM agents, and automated reconciliation systems detecting operational exceptions, billing leakages, and contract discrepancies. 📊 Dataset Summary Total Records: 10,000 standardized enterprise transactions Domains Covered: Logistics & Courier Freight… See the full description on the dataset page: https://huggingface.co/datasets/priteshloke/enterprise-operations-benchmark.texttabular-classification10K<n<100K0 likes54 downloads1mo agoHugging FaceKarmane /enterprise-rag-internal-knowledge-search-benchmark-sample Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.tabulartext-generationn<1K2 likes23 downloads4mo agoHugging FaceKarmane /enterprise-rag-internal-knowledge-search-benchmarkgated Enterprise RAG and Internal Knowledge Search Benchmark Dataset This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.tabulartext-generationn<1K0 likes9 downloads4mo agoHugging Face