CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Krichna /SceneText150kimage100K<n<1M0 likes4.6k downloads2y agoHugging Face02Krispin /ffp FoundationPose paired synthetic renders Each scene row contains two synchronized views and per-object pose, mask ID, bounding box, visibility, and occlusion annotations. Depth is stored as the original float32 NPY bytes; RGB and uint32 instance masks are stored as PNG bytes. Matrices use row-major flattened arrays and canonical column-vector names (camera_from_world and world_from_object). Raw source transforms are retained alongside them. The assets tables contain stable source… See the full description on the dataset page: https://huggingface.co/datasets/Krispin/ffp.tabular100K<n<1M0 likes3.5k downloads3mo agoHugging Face03krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face04KristianS7 /prepacked-fineweb-edu-llama2-32K-T2048 prepacked-fineweb-edu-llama2-32K-T2048 Pre-tokenized and BOS-aligned best-fit packed version of FineWeb-Edu for training with looped nanochat. Tokenized with the Llama 2 tokenizer (32,000 base vocab + 8 special tokens = 32,008). Stats Train split Source karpathy/fineweb-edu-100b-shuffle (1,821 shards) Total tokens 63.26B Total docs 97.1M Rows 30,873,598 Shards 2,059 (train-00000 to train-02058) Rows per shard ~15,000… See the full description on the dataset page: https://huggingface.co/datasets/KristianS7/prepacked-fineweb-edu-llama2-32K-T2048.text-generation0 likes1.6k downloads6mo agoHugging Face05krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face06Abhay786 /krishna_wallpapersimage10K<n<100K0 likes1.3k downloads12d agoHugging Face07krishnakalyan3 /emo_webdsaudio10K<n<100K5 likes1.3k downloads2y agoHugging Face08krishnateja95 /ImageNet-Think ImageNet-Think 250K ImageNet-Think 250K is a large-scale synthetic multimodal reasoning dataset containing of 250,000 images sampled from ImageNet-21K dataset. For each image, we provide a prompt and two different step-by-step reasoning tokens and outputs (answers), enabling evaluation and training for Vision Language Models on reasoning tasks. This dataset is primarily designed for research on multimodal summarization. Installation & Setup Before downloading… See the full description on the dataset page: https://huggingface.co/datasets/krishnateja95/ImageNet-Think.imagesummarization100K<n<1M4 likes1.1k downloads1y agoHugging Face09krisfu /awesome-llm-datasets-only-Chinesetext18 likes874 downloads3y agoHugging Face10kriztahimic /sae-code-correctness-dataimagen<1K0 likes619 downloads6mo agoHugging Face11KrishPatel0111 /AgentNet OpenCUA: Open Foundations for Computer-Use Agents 🌐 Website 🔎 Data Viewer 📝 Paper 💻 Code AgentNet Dataset AgentNet is the first large-scale desktop computer-use agent trajectory dataset, containing 22.6K human-annotated computer-use tasks across Windows, macOS, and Ubuntu systems. Applications This dataset enables training and evaluation of: Vision-language-action (VLA) models for computer use Multi-modal agents… See the full description on the dataset page: https://huggingface.co/datasets/KrishPatel0111/AgentNet.image-text-to-text0 likes554 downloads8mo agoHugging Face12Krishna5T /trial_dataset Trial Dataset (VQA) This dataset contains various configurations for Visual Question Answering (VQA) tasks involving tables, figures, and multiple-choice options. Dataset Structure The dataset is split into multiple configurations based on the complexity of the input (number of tables/figures) and the response type (MCQ or Constructed Response). How to Load You can load any specific configuration using the datasets library: from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/Krishna5T/trial_dataset.imagevisual-question-answeringn<1K0 likes535 downloads6mo agoHugging Face13RaiyanKhaan /KrishokChat KrishokChat Dataset KrishokChat is a provenance-traceable, multi-task Bengali agricultural dataset for safety-critical chemical advisory and domain-specific natural language understanding. Every instance in the dataset retains citation-level provenance (publisher, document title, page range, section path) linked directly to official agricultural extension handbooks and research manuals issued by government and NGO agricultural institutions in Bangladesh. Figure 1:… See the full description on the dataset page: https://huggingface.co/datasets/RaiyanKhaan/KrishokChat.tabular100K<n<1M1 likes515 downloads2mo agoHugging Face14kristina-shemet /train_dataset_21_12text1K<n<10K0 likes483 downloads2y agoHugging Face15open-law-data-thailand /ocs-krisdika Open Law Data Thailand: OCS Krisdika Dataset ชุดข้อมูลกฎหมายจาก สำนักงานคณะกรรมการกฤษฎีกา (Office of the Council of State) รวบรวมและจัดทำโดยโครงการ Open Law Data Thailand เพื่อส่งเสริมการเข้าถึงข้อมูลกฎหมายในรูปแบบที่เครื่องอ่านได้ (Machine-Readable) Dataset Structure ข้อมูลถูกจัดเก็บในรูปแบบ JSON Lines (.jsonl) แบ่งไฟล์ตาม ปีและเดือน (YYYY/YYYY-MM.jsonl) เพื่อความสะดวกในการดาวน์โหลดและบริหารจัดการ Data Fields แต่ละบรรทัด (Row) ประกอบด้วยข้อมูลดังนี้: title… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/ocs-krisdika.text-retrieval4 likes479 downloads10mo agoHugging Face16kriyam-ai /kriyam-tamperflow Kriyam TamperFlow The first document tampering detection benchmark built specifically for Indian documents, with a built-in compression stress-test that exposes how quickly forensic models degrade on real-world scanned material. Dataset Summary State-of-the-art document forgery detectors — CAT-Net, DTD, MVSS-Net, CAFTB, and others — rely on JPEG compression artifacts as their primary forensic signal: inconsistencies in DCT coefficients, block boundaries, and… See the full description on the dataset page: https://huggingface.co/datasets/kriyam-ai/kriyam-tamperflow.imageimage-classification1K<n<10K0 likes465 downloads20d agoHugging Face17krisbailey /cosmopedia-10B Cosmopedia 10B Dataset Description This is a 10.53 Billion token subset of the HuggingFaceTB/cosmopedia dataset. It was created by sampling approximately 45% of each subset (web_samples, stories, stanford, etc.) from the original dataset and deduplicating to ensure high utility. Motivation The original Cosmopedia dataset is massive (~25B+ tokens) and high quality. This 10B version serves as a "Goldilocks" dataset—large enough for meaningful pre-training… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-10B.texttext-generation10M<n<100M0 likes445 downloads8mo agoHugging Face18krishnakalyan3 /kolors-20kimage10K<n<100K5 likes440 downloads2y agoHugging Face19KrisMinchev /finemath-4plus-tokenizedtabular1M<n<10M0 likes411 downloads9mo agoHugging Face20krishnapal2308 /SROIEimage10K<n<100K0 likes410 downloads2y agoHugging Face21Krishnam2004 /recursion-cellular-image-classification-datasetimagen<1K0 likes365 downloads10mo agoHugging Face22krisbailey /RedPajama-Data-V2-1B RedPajama-Data-V2 1B Dataset Description This is a 1.01 Billion token subset of the togethercomputer/RedPajama-Data-V2 dataset (specifically derived from the sample-10B config). It was created by randomly sampling the source data. Motivation RedPajama V2 is a state-of-the-art web dataset with rich quality signals. This 1B token subset allows for rapid testing of these quality signals or other filtering experiments without needing to process the full… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-Data-V2-1B.texttext-generation100K<n<1M0 likes346 downloads8mo agoHugging Face23krishna112003 /Spotify_Music_Analytics_and_Popularity_Prediction0 likes310 downloads4mo agoHugging Face24krisbailey /RedPajama-10B-Weighted RedPajama-10B-Weighted A canonical 10 Billion token weighted subset of the RedPajama-Data-1T dataset. Dataset Description This dataset is a faithful reproduction of the original RedPajama-Data-1T distribution, scaled down to exactly 10 Billion tokens. It is designed to preserve the exact domain ratios of the original dataset (excluding the defunct 'Books' subset). This allows researchers and developers to prototype, debug, and test on a representative slice of the data… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/RedPajama-10B-Weighted.texttext-generation1M<n<10M0 likes286 downloads9mo agoHugging Face25Krinos /MSU-Bench MSU-Bench: Musical Score Understanding Benchmark Evaluating Large Language Models' Comprehension of Complete Musical Scores Overview MSU-Bench is a human-curated benchmark for evaluating the musical score understanding capabilities of Large Language Models (LLMs) and Vision-Language Models (VLMs). It supports multimodal evaluation through both textual (ABC notation) and visual (PDF/image) inputs. Key Statistics: 150 complete musical scores 1,800 generative… See the full description on the dataset page: https://huggingface.co/datasets/Krinos/MSU-Bench.imagequestion-answering1K<n<10K1 likes278 downloads2mo agoHugging Face26krisbailey /cosmopedia-1b Cosmopedia 1B Dataset Description This is a 1 Billion token subset of the krisbailey/cosmopedia-10B dataset, which itself is a 10B subset of HuggingFaceTB/cosmopedia. It was created by uniformly sampling approximately 9.5% of the 10B dataset, ensuring the data distribution remains consistent with the source. Motivation While the 10B dataset is a "Goldilocks" size for many experiments, 1B tokens is the standard size for rapid prototyping, scaling law… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/cosmopedia-1b.texttext-generation1M<n<10M0 likes273 downloads8mo agoHugging Face27KrisMinchev /generated-finemath-292968-node-2text100K<n<1M0 likes264 downloads8mo agoHugging Face28krisbailey /falcon-refinedweb-1B Falcon RefinedWeb 1B Dataset Description This is a 1.01 Billion token subset of the tiiuae/falcon-refinedweb dataset. It was created by streaming the dataset with a large shuffle buffer to ensure a random, representative sample of the web data. Motivation RefinedWeb is a high-quality filtered web dataset, but the full version is massive. This 1B token slice provides a perfect testbed for evaluating model architecture changes or for use in curriculum learning… See the full description on the dataset page: https://huggingface.co/datasets/krisbailey/falcon-refinedweb-1B.texttext-generation1M<n<10M0 likes254 downloads8mo agoHugging Face29kritsadaK /EDGAR-CORPUS-Financial-Summarization EDGAR-CORPUS : 10K Financial Report Summarization Extracted from SEC EDGAR filings (1993-2020). This dataset enhances financial report summarization by leveraging a hybrid AI model strategy. Using: ChatGPT-3.5 Turbo(~70%), Claude 3.5 (~30% to generate structured, accurate, and concise summaries) Dataset Composition Summaries in this dataset are generated using a hybrid AI model strategy, balancing quality and efficiency:ChatGPT-3.5 Turbo (~70%) – Used for structured… See the full description on the dataset page: https://huggingface.co/datasets/kritsadaK/EDGAR-CORPUS-Financial-Summarization.textsummarization10K<n<100K4 likes253 downloads2y agoHugging Face30KrisMinchev /generated-finemath-585942-node-0text100K<n<1M0 likes253 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.