datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
layergen-eval-latents
LayerGen — Eval-Set Latents (VAE latents + baked text embeddings)
Pre-encoded evaluation-set inputs for the LayerGen layer-decomposition / harmonization models,
so inference can run anywhere (off-AIP) without the raw video → VAE-encode → umT5-encode pipeline.
Each *.parquet is one clip and is fully self-contained:
column group
contents
{composite,mask,fg,bg}_latent_bytes (+ _shape, _dtype)
4-stream Wan-VAE latents, 81f/21 latent-T, fp16, [16,21,60,104]… See the full description on the dataset page: https://huggingface.co/datasets/cs-mshah/layergen-eval-latents.csmar-legacycsmd
Dataset Card for "Continuous Scale Meaning Dataset" (CSMD)
CSMD was created for MeaningBERT: Assessing Meaning Preservation Between Sentences.
It contains 1,355 English text simplification meaning preservation annotations. Meaning preservation measures how well the meaning of the output text corresponds to the meaning of the source (Saggion, 2017).
The annotations were taken from the following four datasets:
ASSET
QuestEVal,
SimpDa_2022 and,
Simplicity-DA.
It contains a data… See the full description on the dataset page: https://huggingface.co/datasets/graalul/csmd.Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning
Curated Arabic Speech Dataset for Seasme (from MCV17)
Dataset Description
This dataset is a curated and preprocessed version of the Arabic (ar) subset from Mozilla Common Voice (MCV) 17.0. It has been specifically prepared for fine-tuning conversational speech models, with a primary focus on the Seasme-CSM model architecture. The dataset consists of audio clips in WAV format (24kHz, mono) and their corresponding transcripts, along with integer speaker IDs.
The original… See the full description on the dataset page: https://huggingface.co/datasets/MAdel121/Common-Voice-17-Arabic-for-Seasme-CSM-Finetuning.GPTscience_maths_csmlmodel_df_patentSBERTacs_mall-product-reviews
Dataset Card for Mall.cz Product Reviews (Czech)
Dataset Description
The dataset contains user reviews from Czech eshop <mall.cz>
Each review contains text, sentiment (positive/negative/neutral), and automatically-detected language (mostly Czech, occasionaly Slovak) using lingua-py
The dataset has in total (train+validation+test) 30,000 reviews. The data is balanced.
Train set has 8000 positive, 8000 neutral and 8000 negative reviews.
Validation and test set each have… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_mall-product-reviews.CSMBench
CSMBench
Benchmarking Cross-Scale Perception Ability of Large Multimodal Models in Material Science.
Config(在 Config/Subset 下拉框中选择)
multi_scale: Open-ended questions / Image matching (image_caption, describe_paragraph)
multi_scale_mcq: Multiple-choice questions (options, correct_answer)
how-people-make-money-csm1bbuzz_sources_026_GPTscience_maths_csmlconversations_datasetMaterials-Informatics
Dataset Card for "Materials-Informatics"
Dataset Name: Materials-Informatics
Dataset Owner: cs-mubashir
Language: English
Size: ~600+ entries
Last Updated: May 2025
Source: Extracted from arxiv dataset research repository
Dataset Summary
The Materials-Informatics dataset is a curated collection of research papers from arxiv repository focusing on the intersection
of artificial intelligence (AI) and materials science and engineering (MSE). Each entry provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/cs-mubashir/Materials-Informatics.csml-debug-v0CSMC_2016_NXSR_Commercial_Mortgage_Trust_1691198
cik
form
accessionNumber
fileNumber
filmNumber
reportDate
url
1691198
ABS-EE
0001539497-16-004254
333-207361-04
162039674
2016-12-07
https://sec.gov/Archives/edgar/data/1691198/000153949716004254
1691198
ABS-EE
0001056404-17-000231
333-207361-04
17563185
2017-02-01
https://sec.gov/Archives/edgar/data/1691198/000105640417000231
1691198
ABS-EE
0001056404-17-000574
333-207361-04
17653893
2017-02-17
https://sec.gov/Archives/edgar/data/1691198/000105640417000574
1691198
ABS-EE… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/CSMC_2016_NXSR_Commercial_Mortgage_Trust_1691198.pymatgen-github-issuesfinetuned-lb-ar-csm-3-5h-groupedcsm-turkish-ttsadobe_csm_finetuneCSMVS-Museum-Img-QAfinetuned-lb-ar-csm-3-1h-groupedparquet-file
CSMAR Parquet 数据集
本仓库包含大量彼此独立、字段结构不同的 CSMAR 业务表。
Dataset Viewer 默认展示轻量的数据表目录 viewer/table_catalog.parquet;完整业务数据请在 Files and versions 中按系列、数据库和表路径访问对应 Parquet 文件。
llama3_testmorgan-csmfinetuned-lb-ar-csm-3-2h-groupedfinetuned-lb-ar-csm-3-full-groupedpartial_GH_cleaned_trimmedza-zenande-csm-metadata2-v1us-julia-csm-v1ccnews_www.csmonitor_scmfinetuned-lb-ar-csm-3-3h-grouped
