CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RaiyanKhaan /KrishokChat KrishokChat Dataset KrishokChat is a provenance-traceable, multi-task Bengali agricultural dataset for safety-critical chemical advisory and domain-specific natural language understanding. Every instance in the dataset retains citation-level provenance (publisher, document title, page range, section path) linked directly to official agricultural extension handbooks and research manuals issued by government and NGO agricultural institutions in Bangladesh. Figure 1:… See the full description on the dataset page: https://huggingface.co/datasets/RaiyanKhaan/KrishokChat.tabular100K<n<1M2 likes478 downloads2mo agoHugging Face02nyu-dice-lab /lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private Dataset Card for Evaluation run of Kukedlc/Neural-Krishna-Multiverse-7b Dataset automatically created during the evaluation run of model Kukedlc/Neural-Krishna-Multiverse-7b The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-private.tabular100K<n<1M0 likes184 downloads2y agoHugging Face03hari-krishna-ai /enterprise-text-to-sql-benchmark Enterprise Text-to-SQL Benchmark 3,087 natural-language questions paired with executable PostgreSQL, over a 12-table enterprise schema (sales, catalogue, logistics, HR). Built to answer one question honestly: does fine-tuning actually improve text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict execution accuracy from 43.71 % to 68.43 %, and 70.86 % with a self-correction loop — and the benchmark is designed so that number cannot be inflated by leakage or by… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.texttable-question-answering1K<n<10K0 likes129 downloads2d agoHugging Face04KritishBokde91 /racf-ivit-finaltabularn<1K0 likes124 downloads25d agoHugging Face05krishnakartik /gemma4-social-bias-judge-pairs gemma4-social-bias-judge-pairs Training and evaluation data for the judge-from-scratch project, which fine-tuned Gemma 4 E4B into a specialist social-bias judge (primary model, SFT-only secondary). This dataset contains: sft.jsonl (3,844 rows) — the SFT training set, in TRL prompt-completion shape. 1,922 base pairs surviving the post-label confidence filter (15 low-confidence rows dropped from the 1,938-pair labeling input), doubled by position swap to teach the judge to mirror… See the full description on the dataset page: https://huggingface.co/datasets/krishnakartik/gemma4-social-bias-judge-pairs.texttext-classification10K<n<100K0 likes99 downloads5mo agoHugging Face06krittus /k12-indian-curriculum-4.9m BharatLLM K-12 Indian Curriculum Dataset (4.9M) 4,904,936 question-answer pairs covering CBSE/NCERT K-12 curriculum across 12 Indian languages. Language Script Entries English Latin ~594K Hindi Devanagari ~449K Bengali Bengali ~408K Telugu Telugu ~408K Tamil Tamil ~408K Kannada Kannada ~408K Malayalam Malayalam ~408K Marathi Devanagari ~408K Gujarati Gujarati ~408K Odia Odia ~408K Punjabi Gurmukhi ~408K Urdu Nastaliq ~374K Format… See the full description on the dataset page: https://huggingface.co/datasets/krittus/k12-indian-curriculum-4.9m.textquestion-answering1M<n<10M1 likes98 downloads6mo agoHugging Face07KrisQ /utilmemtextn<1K0 likes82 downloads27d agoHugging Face08KrisQ /StudyChat StudyChat (Public Mirror) Public ungated mirror of wmcnicho/StudyChat. tabular10K<n<100K0 likes81 downloads6mo agoHugging Face09krisztiankoos /uldr-v0.1-pilot Ukrainian Language Decolonization & Reasoning (ULDR) — Pilot Canary Release v0.1 [!IMPORTANT] Exploratory Pilot / Canary Release (v0.1): This dataset represents an early exploratory pilot canary release (v0.1-pilot) establishing our baseline data pipeline, schema contracts, and directional validation. It is not the final production release (v1.0). The full production release is scheduled for Phase 5.5 following complete evaluation suite assembly, dialect & historical protection… See the full description on the dataset page: https://huggingface.co/datasets/krisztiankoos/uldr-v0.1-pilot.texttext-generation10K<n<100K0 likes80 downloads10d agoHugging Face10krishgoel /chronocept Chronocept: Instilling a Sense of Time in Machines Authors: Krish Goel, Sanskar Pandey, KS Mahadevan, Harsh Kumar, and Vishesh KhadariaPublication: Chronocept: Instilling a Sense of Time in Machines Chronocept is a benchmark for modeling the temporal validity of textual information as a continuous probability distribution over time. By fitting skewed-normal curves to annotated facts and passages, Chronocept captures phenomena such as gradual decay, delayed onset, and asymmetric peak… See the full description on the dataset page: https://huggingface.co/datasets/krishgoel/chronocept.textfeature-extraction1K<n<10K3 likes71 downloads1y agoHugging Face11nyu-dice-lab /lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-v3-private Dataset Card for Evaluation run of Kukedlc/Neural-Krishna-Multiverse-7b-v3 Dataset automatically created during the evaluation run of model Kukedlc/Neural-Krishna-Multiverse-7b-v3 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Kukedlc-Neural-Krishna-Multiverse-7b-v3-private.tabular100K<n<1M0 likes68 downloads2y agoHugging Face12aiuser3993 /Llama-Krikri-8B-Instruct-ChatCreative and vocabulary-rich (LLM-focused) Greek conversation data generated by ilsp/Llama-Krikri-8B-Instruct-GGUF for distlization purposes. textn<1K0 likes63 downloads28d agoHugging Face13hari-krishna-ai /text-to-sql-eval-predictions What the text-to-SQL models actually generated Every prediction behind the numbers in qwen3-8b-text2sql-qlora: the 453 test questions of the enterprise text-to-SQL benchmark, each answered by four configurations of the same model, each answer executed against the reference PostgreSQL database and scored by comparing result sets. 1,812 rows. I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.tabulartext-generation1K<n<10K0 likes59 downloads8d agoHugging Face14hari-krishna-ai /text-to-sql-phrasing-robustness Does sloppy phrasing break text-to-SQL? The enterprise text-to-SQL benchmark lists its own biggest caveat: every question is template-generated, so real user phrasing is untested. This is the test. 35 test questions (one per template), each sent to the deployed pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment. 24 questions and 85 answers survive the filter described under Setup; every answer was executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.tabulartext-generationn<1K0 likes54 downloads7d agoHugging Face15KrithikV /MedDistractQAtext1K<n<10K0 likes51 downloads1y agoHugging Face16krishan90 /Onion-Classification-and-Segmentation-Dataset Onion Classification and Segmentation Dataset The current agricultural industry faces challenges of high cost and low efficiency in crop monitoring and pest and disease detection. Existing solutions often rely on manual inspection, which is inefficient and prone to errors. This dataset aims to enhance the application capability of computer vision models in agriculture by providing high-quality onion images and segmentation annotations. The dataset construction process includes… See the full description on the dataset page: https://huggingface.co/datasets/krishan90/Onion-Classification-and-Segmentation-Dataset.textimage-segmentationn<1K0 likes45 downloads7d agoHugging Face17krisfu /data_agent_workflow_generate_medical_datacreate by data agent workflow text10K<n<100K1 likes44 downloads1y agoHugging Face18krisfu /delicate_medical_r1_data_chinese 数据: 基于华佗开源的高质量语料库 技术链路: 多智能体 + 数据进化 + 推理过程生成 + 推理过程验证过滤 多智能体 利用metagpt搭建的造数据的workflow 数据进化 改造 self-instruct 进行数据进化 推理过程生成 使用 qwq 模型进行每 query 10次 think 过程的生成 推理过程验证过滤所使用的指标 质量评估 2.1.1. 召回率 首先经过step_partition步骤后,得到了带有N步推理的思维链。使用llm找到其中的关键 步骤,然后依次判断这些关键步骤解决的问题或陈述的事实是否在出现在真实答案中。将出现的个数除以总个数, 即可得到召回率分值。 2.1.2. 精确率 首先经过step_partition步骤后,得到了带有N步推理的思维链。使用llm,以真实答案为… See the full description on the dataset page: https://huggingface.co/datasets/krisfu/delicate_medical_r1_data_chinese.text1K<n<10K4 likes43 downloads1y agoHugging Face19KritiAI /Xijinping-TTS-Voicebank 习近平音源 所有声音资料来自公开影像,属于公有领域目前有 1h30m 的截取后声音,足够进行 Fine-tuning Usage 按句截取 python -m pip install -r requirement.txt python split.py 新增声音资料后,使用 Whisper 产生带有时间标记的 JSON 档,并手动复制到 ./voice/[FILE].json export OPENAI_API_KEY="API_KEY_HERE" python whisper.py ./[FILE].[AUDIO_EXTENSION] 产生 Bert-VITS2 微调所需的 esd.list 档案 python index_to_list.py audiotext-to-speech10K<n<100K4 likes41 downloads1y agoHugging Face20krishy-d /formatbench FormatBench: A Preference Dataset for Correcting LLM Formatting Bias Part of the Prosify project. Also available on Kaggle. The problem this dataset addresses Large language models trained with RLHF systematically over-format their outputs are defaulting to bullet points, bold headers, and templated structures even when flowing prose would serve the reader better. This shows up most visibly when people use LLMs for real-world tasks: a request to "polish this… See the full description on the dataset page: https://huggingface.co/datasets/krishy-d/formatbench.texttext-generationn<1K0 likes41 downloads4mo agoHugging Face21kristofferkjeldby /KjeldChat-1.1B-rag-index KjeldChat 1.1B — RAG passage index The retrieval half of KjeldChat 1.1B: a prebuilt passage index over Wikipedia, so the model can answer from a retrieved passage instead of from a 1.1B model's thin parametric memory. Published so the Space can load it at startup, and so anyone running the model locally doesn't have to spend hours refetching and re-embedding it. Contents 45,104 passages, embedded with all-MiniLM-L6-v2. File Size What it is… See the full description on the dataset page: https://huggingface.co/datasets/kristofferkjeldby/KjeldChat-1.1B-rag-index.tabularquestion-answeringn<1K0 likes41 downloads2mo agoHugging Face22krishnareddy /triage-questions Medical Triage Complaint Data Structure README This data structure is designed to use for supervised finetuining of llam2 over generating triage questions based on provided patient complaint/age/gender as input JSON Format The data structure is represented in JSON format, with two main sections: input and questions. Input Section The input section contains information about the patient's complaint, age, and gender. { "input": { "complaint": "Patient's… See the full description on the dataset page: https://huggingface.co/datasets/krishnareddy/triage-questions.textn<1K3 likes35 downloads3y agoHugging Face23shjkv /krispintech-federal-source-index Airfoil Aerodynamics & Vortex Method Reference Compiled by Krispin Technologies Archive — https://krispintech.com 14 sections distilled from 14 public-domain United States federal publications (ntrs.nasa.gov). Every figure, limit and clause number is reproduced unchanged; each record carries the original document URL so any value can be checked at the source. No copyrighted standards are reproduced. Files krispintech-source-index.csv — lookup table, one row per… See the full description on the dataset page: https://huggingface.co/datasets/shjkv/krispintech-federal-source-index.textn<1K0 likes33 downloads6d agoHugging Face24alledged /kriptik-uicoder-trainingtext1K<n<10K1 likes28 downloads8mo agoHugging Face25Krishnapadala55 /brahmastra-benchmark BRAHMASTRA Security LLM Benchmark Suite A 6-suite, 280-prompt benchmark for evaluating Large Language Models on Web Application Security Testing (DAST) tasks. This benchmark accompanies the release of BRAHMASTRA v0.3 and provides a reproducible methodology for measuring DAST-relevant capabilities of security-fine-tuned LLMs. Why this benchmark? Existing security LLM benchmarks (CyberSecEval, SecQA, HackBench) focus on penetration-testing scenarios or general security… See the full description on the dataset page: https://huggingface.co/datasets/Krishnapadala55/brahmastra-benchmark.texttext-classificationn<1K0 likes24 downloads5mo agoHugging Face26KrisPi /PythonTutor-Evol-1k-DPO-GPT4_vs_35Started with: https://huggingface.co/datasets/nickrosh/Evol-Instruct-Code-80k-v1 (GPT-3.5 Turbo) Randomly selected 1000 where output contained "```python" in output Generated GPT-4 answers to those for the sake of LIMA-like "Python Tutor" Instruct fine-tuning as well as validate DPO Fine-Tuning (where GPT-4 answers will be preferred to GPT-3.5 Turbo) Then filtered refusals (looking for "impossible" or "sorry") GPT-4 System Prompt: You are an intelligent assistant that generates Python code.… See the full description on the dataset page: https://huggingface.co/datasets/KrisPi/PythonTutor-Evol-1k-DPO-GPT4_vs_35.textn<1K14 likes23 downloads3y agoHugging Face27krishna1707 /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/krishna1707/real-toxicity-prompts.tabular10K<n<100K0 likes19 downloads7mo agoHugging Face28krispyATL /pip-onetext1K<n<10K5 likes18 downloads2y agoHugging Face29krispyATL /pip1text1K<n<10K0 likes17 downloads2y agoHugging Face30open-llm-leaderboard /ilsp__Llama-Krikri-8B-Instruct-detailsgated Dataset Card for Evaluation run of ilsp/Llama-Krikri-8B-Instruct Dataset automatically created during the evaluation run of model ilsp/Llama-Krikri-8B-Instruct The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ilsp__Llama-Krikri-8B-Instruct-details.tabular10K<n<100K0 likes17 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.