CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cs-mshah /layergen-eval-latents LayerGen — Eval-Set Latents (VAE latents + baked text embeddings) Pre-encoded evaluation-set inputs for the LayerGen layer-decomposition / harmonization models, so inference can run anywhere (off-AIP) without the raw video → VAE-encode → umT5-encode pipeline. Each *.parquet is one clip and is fully self-contained: column group contents {composite,mask,fg,bg}_latent_bytes (+ _shape, _dtype) 4-stream Wan-VAE latents, 81f/21 latent-T, fp16, [16,21,60,104]… See the full description on the dataset page: https://huggingface.co/datasets/cs-mshah/layergen-eval-latents.tabularvideo-to-video1K<n<10K0 likes1.1k downloads1mo agoHugging Face02cs-mshah /layergen-evalsvideon<1K0 likes593 downloads1mo agoHugging Face03cs-mshah /layergen-evals-baselines0 likes491 downloads1mo agoHugging Face04Zion-HF /csmar-legacytext2 likes368 downloads1y agoHugging Face05graalul /csmd Dataset Card for "Continuous Scale Meaning Dataset" (CSMD) CSMD was created for MeaningBERT: Assessing Meaning Preservation Between Sentences. It contains 1,355 English text simplification meaning preservation annotations. Meaning preservation measures how well the meaning of the output text corresponds to the meaning of the source (Saggion, 2017). The annotations were taken from the following four datasets: ASSET QuestEVal, SimpDa_2022 and, Simplicity-DA. It contains a data… See the full description on the dataset page: https://huggingface.co/datasets/graalul/csmd.texttext-classification1K<n<10K2 likes318 downloads3y agoHugging Face06cs-mshah /SynMirror Dataset Card for SynMirror This repository hosts the data for Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror Reflections (accepted at 3DV'25).SynMirror is a first-of-its-kind large scale synthetic dataset on mirror reflections, with diverse mirror types, objects, camera poses, HDRI backgrounds and floor textures. Dataset Details Dataset Description SynMirror consists of samples rendered from 3D assets of two widely used 3D… See the full description on the dataset page: https://huggingface.co/datasets/cs-mshah/SynMirror.text-to-image100K<n<1M1 likes217 downloads1y agoHugging Face07giannisan /gemma-csm-shards gemma-csm curated shards Encoded training shards for the gemma-csm project (Mimi RVQ codes + text). Derived from public corpora: LJSpeech (public domain), LibriTTS (CC BY 4.0), OpenS2S / CASIA-LM. These are compressed code representations, not raw audio. shards : LJSpeech, single voice (tag [lj]), short clips shards_long_lj : LJSpeech concatenated to ~16-20s long-format segments shards_libritts / shards_libritts_o : LibriTTS multispeaker shards_long :… See the full description on the dataset page: https://huggingface.co/datasets/giannisan/gemma-csm-shards.0 likes195 downloads3mo agoHugging Face08csmet /xtc_pillsimagen<1K0 likes78 downloads3y agoHugging Face09MAdel121 /Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning Curated Arabic Speech Dataset for Seasme (from MCV17) Dataset Description This dataset is a curated and preprocessed version of the Arabic (ar) subset from Mozilla Common Voice (MCV) 17.0. It has been specifically prepared for fine-tuning conversational speech models, with a primary focus on the Seasme-CSM model architecture. The dataset consists of audio clips in WAV format (24kHz, mono) and their corresponding transcripts, along with integer speaker IDs. The original… See the full description on the dataset page: https://huggingface.co/datasets/MAdel121/Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning.audio10K<n<100K1 likes66 downloads1y agoHugging Face10knowrohit07 /GPTscience_maths_csmltext100K<n<1M5 likes49 downloads3y agoHugging Face11csmoilis /model_df_patentSBERTatabular10K<n<100K0 likes42 downloads7mo agoHugging Face12fewshot-goes-multilingual /cs_mall-product-reviews Dataset Card for Mall.cz Product Reviews (Czech) Dataset Description The dataset contains user reviews from Czech eshop <mall.cz> Each review contains text, sentiment (positive/negative/neutral), and automatically-detected language (mostly Czech, occasionaly Slovak) using lingua-py The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced. Train set has 8000 positive, 8000 neutral and 8000 negative reviews. Validation and test set each have… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_mall-product-reviews.texttext-classification10K<n<100K1 likes39 downloads4y agoHugging Face13lututu /CSMBench CSMBench Benchmarking Cross-Scale Perception Ability of Large Multimodal Models in Material Science. Config(在 Config/Subset 下拉框中选择) multi_scale: Open-ended questions / Image matching (image_caption, describe_paragraph) multi_scale_mcq: Multiple-choice questions (options, correct_answer) imageimage-to-text1K<n<10K0 likes35 downloads6mo agoHugging Face14CSML-IIT /encoderops0 likes29 downloads1y agoHugging Face15suyashkrishangarg /how-people-make-money-csm1baudion<1K0 likes21 downloads3mo agoHugging Face16supergoose /buzz_sources_026_GPTscience_maths_csmltext100K<n<1M0 likes18 downloads2y agoHugging Face17Codyfederer /csm-podcast-tokenized10K<n<100K0 likes16 downloads8mo agoHugging Face18csmcvnc /conversations_datasettext1K<n<10K0 likes13 downloads2y agoHugging Face19cs-mubashir /Materials-Informatics Dataset Card for "Materials-Informatics" Dataset Name: Materials-Informatics Dataset Owner: cs-mubashir Language: English Size: ~600+ entries Last Updated: May 2025 Source: Extracted from arxiv dataset research repository Dataset Summary The Materials-Informatics dataset is a curated collection of research papers from arxiv repository focusing on the intersection of artificial intelligence (AI) and materials science and engineering (MSE). Each entry provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/cs-mubashir/Materials-Informatics.textn<1K0 likes13 downloads1y agoHugging Face20kejian /csml-debug-v0textn<1K0 likes10 downloads3y agoHugging Face21DenyTranDFW /CSMC_2016_NXSR_Commercial_Mortgage_Trust_1691198 cik form accessionNumber fileNumber filmNumber reportDate url 1691198 ABS-EE 0001539497-16-004254 333-207361-04 162039674 2016-12-07 https://sec.gov/Archives/edgar/data/1691198/000153949716004254 1691198 ABS-EE 0001056404-17-000231 333-207361-04 17563185 2017-02-01 https://sec.gov/Archives/edgar/data/1691198/000105640417000231 1691198 ABS-EE 0001056404-17-000574 333-207361-04 17653893 2017-02-17 https://sec.gov/Archives/edgar/data/1691198/000105640417000574 1691198 ABS-EE… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/CSMC_2016_NXSR_Commercial_Mortgage_Trust_1691198.tabular10K<n<100K0 likes9 downloads2y agoHugging Face22jackynix /CSMV_visualThis repository contains the visual features of the CSMV dataset released in Paper Infer Induced Sentiment of Comment Response to Video: A New Task, Dataset and Baseline. The repository contains feature representations of the micro-videos. Each subfolder is named after a different feature extraction method, and the features for each video are saved as .npy files. The filenames correspond to the video_file_id. Currently, features extracted using I3D(recommend) and R(2+1)D have been released.… See the full description on the dataset page: https://huggingface.co/datasets/jackynix/CSMV_visual.video10K<n<100K1 likes9 downloads1y agoHugging Face23cs-mubashir /pymatgen-github-issuestabular1K<n<10K0 likes7 downloads1y agoHugging Face24ChristopherY10 /finetuned-lb-ar-csm-3-5h-groupedaudio1K<n<10K0 likes7 downloads6mo agoHugging Face25surferrosa22 /csm-turkish-ttsaudio10K<n<100K0 likes7 downloads4mo agoHugging Face26ShreyansJain04 /adobe_csm_finetunetext1K<n<10K0 likes6 downloads3y agoHugging Face27ZacJQ /CSMVS-Museum-Img-QAimage10K<n<100K0 likes6 downloads2y agoHugging Face28ChristopherY10 /finetuned-lb-ar-csm-3-1h-groupedaudio1K<n<10K0 likes5 downloads6mo agoHugging Face29CSML-IIT /symm_learning0 likes5 downloads6mo agoHugging Face30d-csmar /parquet-filegated CSMAR Parquet 数据集 本仓库包含大量彼此独立、字段结构不同的 CSMAR 业务表。 Dataset Viewer 默认展示轻量的数据表目录 viewer/table_catalog.parquet;完整业务数据请在 Files and versions 中按系列、数据库和表路径访问对应 Parquet 文件。 text1K<n<10K1 likes5 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.