CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saidurga001301 /mathmetics-dataset-custom Transformer Math Dataset (54,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 54,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 1 to 2 Integer Operand Ratio: 0% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-custom.texttext-generation100M<n<1B0 likes713 downloads1mo agoHugging Face02saidurga001301 /mathmetics-dataset-intmax Transformer Math Dataset (200,000,000 Samples Sharded) High-precision synthetic mathematical expression dataset generated for training sequence-to-sequence math Transformers in JAX/Flax. Dataset Structure Total Samples: 200,000,000 Shard Format: JSONL sharded files (100,000 samples per shard) Supported Operations: +, -, *, /, ^, sin, cos, tan, log, ln, exp, sqrt, abs Expression Depth Range: Depth 4 to 6 Integer Operand Ratio: 80% Data Fields Each… See the full description on the dataset page: https://huggingface.co/datasets/saidurga001301/mathmetics-dataset-intmax.texttext-generation100M<n<1B0 likes545 downloads1mo agoHugging Face03saidsef /tech-docs Technical Documentation Dataset A curated collection of technical documentation and guides spanning various cloud-native technologies, infrastructure tools, and machine learning frameworks. This dataset contains 1,397 documents in JSONL format, covering essential topics for modern software development and DevOps practices. Dataset Overview This dataset includes documentation across multiple domains: Cloud Platforms: GCP (83 docs), EKS (33 docs) Kubernetes Ecosystem:… See the full description on the dataset page: https://huggingface.co/datasets/saidsef/tech-docs.textquestion-answering1K<n<10K2 likes269 downloads2y agoHugging Face04saidutta69 /Odia-Web-Corpus-v1 Odia Web Corpus v1 The inaugural release of a curated Odia (Oriya) web text corpus. Scraped and filtered from publicly accessible web sources to support Odia language modeling and NLP research. Dataset Details Language: Odia (Oriya, ISO 639-3: ory) Format: JSONL (one JSON object per line) Size: ~650K documents, ~0.9 GB text License: CC-BY-4.0 Data Fields Field Type Description text string Cleaned document body title string… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/Odia-Web-Corpus-v1.texttext-generation100K<n<1M0 likes131 downloads13d agoHugging Face05saidutta69 /red-pill-drug-discovery-formulation 🔴 RED-PILL Research Enhanced Dataset for Pharmaceutical Innovation in Learning & Language The first open instruction-tuning dataset for drug discovery & formulation development. Built for fine-tuning Heretic-ablated models that won't refuse your pharmaceutical R&D questions. ⚡ Quick Start from datasets import load_dataset # Load the full dataset ds = load_dataset("saidutta69/red-pill-drug-discovery-formulation"… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/red-pill-drug-discovery-formulation.texttext-generation1K<n<10K0 likes124 downloads13d agoHugging Face06Sai-1123 /CBT-Bench CBT-Bench Dataset Overview CBT-Bench is a benchmark dataset designed to evaluate the proficiency of Large Language Models (LLMs) in assisting cognitive behavior therapy (CBT). The dataset is organized into three levels, each focusing on different key aspects of CBT, including basic knowledge recitation, cognitive model understanding, and therapeutic response generation. The goal is to assess how well LLMs can support various stages of professional mental health… See the full description on the dataset page: https://huggingface.co/datasets/Sai-1123/CBT-Bench.textquestion-answering1K<n<10K0 likes37 downloads3mo agoHugging Face07SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads4mo agoHugging Face08Sakalti /saishin-abtexttext-generationn<1K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.