datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uploadm8-promo-targeting-v1
UploadM8 promo targeting dataset
Training/evaluation rows exported from UploadM8 (ml_outcome_labels, campaign telemetry).
Populated by admin scripts and optional UM8_HF_SYNC_VISUAL_ENTITIES uploads.
uploadm8-content-success-v1cedr_v1
Dataset Card for [cedr]
Dataset Summary
The Corpus for Emotions Detecting in Russian-language text sentences of different social sources (CEDR) contains 9410 comments labeled for 5 emotion categories (joy, sadness, surprise, fear, and anger).
Here are 2 dataset configurations:
"main" - contains "text", "labels", and "source" features;
"enriched" - includes all "main" features and "sentences".
Dataset with predefined train/test splits.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/sagteam/cedr_v1.agi-biospheric-10-constraints
AGI Biospheric — 10 Core Biospheric Constraints and 52 Interdependencies
Overview
AGI Biospheric is a bilingual (French / English) structured dataset modeling 10 core biospheric constraints and 52 direct interdependencies relevant to long-term reflection on AGI alignment, ecological limits, civilizational resilience, and the material conditions of intelligence.
This repository is the first public machine-readable release of the AGI Biospheric framework. It should… See the full description on the dataset page: https://huggingface.co/datasets/Cedre83/agi-biospheric-10-constraints.TempCloze
TempCloze
Paper: TempCloze: Can Video-LLMs Identify the Missing Middle?
TempCloze is a video cloze benchmark for evaluating whether Video-LLMs can identify the missing middle of a video from its beginning and ending context.
This repository provides metadata for 1,521 videos from seven sources. Each row identifies one source video and the temporal boundaries of its missing segment. The benchmark code and evaluation instructions are available in the official GitHub repository.… See the full description on the dataset page: https://huggingface.co/datasets/CedPei/TempCloze.CEDRClassification
CEDRClassification
An MTEB dataset
Massive Text Embedding Benchmark
Classification of sentences by emotions, labeled into 5 categories (joy, sadness, surprise, fear, and anger).
Task category
t2c
Domains
Web, Social, Blog, Written
Reference
https://www.sciencedirect.com/science/article/pii/S1877050921013247
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/CEDRClassification.cedr-m7
CEDR M7
Russian text emotion corpus, adapted to seven Aniemore classes.
Provenance
This corpus is not ours. It is an adaptation: the source is sagteam/cedr_v1, and Aniemore's contribution is remapping it onto the seven classes the rest of the library uses. Credit for the collection and the original annotation belongs to the CEDR authors.
Splits
Split
Rows
Mean words
Longest
train
7528
14.8
51
test
1882
14.6
43
Fields… See the full description on the dataset page: https://huggingface.co/datasets/Aniemore/cedr-m7.CedPane
CedPane: Chinese-English Dictionary Public-domain Additions for Names Etc
汉英词典公有领域专名等副刊CedPane
From https://ssb22.user.srcf.net/cedpane/
(also mirrored on GitLab Pages just in case)
English summary
People learning Chinese as a foreign language sometimes use software to help them read a text. But when Western names are written using Chinese characters, the result is not always something an average dictionary can help with—the software might give you… See the full description on the dataset page: https://huggingface.co/datasets/real-ssb22/CedPane.data_ClassificationResultscedulaProfecional_Reversocedar-signatures-genuinesignatures-genuine-combined-cedar-beng-hindi
signatures-genuine-combined-cedar-beng-hindi
Mirror of the exact bmcore v24 local holdout subset: 274918 image files.
Benchmark label: real. This preserves the source benchmark label; it is not an independent label review.
Source reference: https://huggingface.co/datasets/indra-inc/signature_genuine_forged_combined_cedar_beng_hindi.
No new license or ownership claim is asserted by this mirror. Original source rights and restrictions remain applicable.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/signatures-genuine-combined-cedar-beng-hindi.cedr-classificationcedulaProfecional_Frontalnumerous-order-1fa8ee
numerous-order-1fa8ee
Synthetic sensors test data: 55 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/cedarloom/numerous-order-1fa8ee.CEDA-215
CEDA-215: Command-Execution Defense Assessment (215 entries)
CEDA-215 is a labeled evaluation dataset for benchmarking input-classification defense pipelines for command-execution LLM chatbots. Each entry pairs a natural-language input with a ground-truth SAFE or UNSAFE label, allowing fine-grained measurement of true positives, false positives, true negatives, and false negatives across an OWASP LLM Top 10 - aligned threat taxonomy.
Why this dataset exists
Existing… See the full description on the dataset page: https://huggingface.co/datasets/aalayed/CEDA-215.CEdit-BenchCEdit-Bench is a comprehensive evaluation suite, first proposed in the LongCat-Image technical report, and developed by integrating and extending existing image editing benchmarks. We further curate new data to enhance task diversity, yielding a robust dataset of 1,464 bilingual (Chinese–English) editing pairs across 15 fine-grained task categories, providing a more holistic and rigorous standard for evaluating image editing models.
An example entry is shown below:
{
"key":… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/CEdit-Bench.similar-grandfather-b33c94
similar-grandfather-b33c94
Synthetic products test data: 43 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/cedarAri/similar-grandfather-b33c94.electrical-protection-99393b
electrical-protection-99393b
Synthetic weather test data: 35 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Steven/electrical-protection-99393b.safe-storage-41a8da
safe-storage-41a8da
Synthetic weather test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/cedarDawnJ/safe-storage-41a8da.cedula_antigua_anverso_v4us-noncompete-enforceability
US Non-Compete Clause Enforceability - Verified Court Holdings
A primary-source-cited dataset of how U.S. courts have actually ruled on employee non-compete (restrictive-covenant) clauses - each outcome quoted verbatim from, and linked to, the source opinion. 76 verified holdings across 30 U.S. states (snapshot 2026-07-05).
Published by Cedarstone Ventures LLC, maker of ClauseDelta. The live, continuously-updated version - with a free clause checker covering all 51 U.S.… See the full description on the dataset page: https://huggingface.co/datasets/cedarstoneventures/us-noncompete-enforceability.aggressive-trouble-61787d
aggressive-trouble-61787d
Synthetic products test data: 46 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-MaryM/aggressive-trouble-61787d.representative-answer-0cb3e0
representative-answer-0cb3e0
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Patricia/representative-answer-0cb3e0.difficult-technology-8dac0d
difficult-technology-8dac0d
Synthetic products test data: 56 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Craft89/difficult-technology-8dac0d.immediate-bend-54d6cd
immediate-bend-54d6cd
Synthetic products test data: 36 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/cedarBridge/immediate-bend-54d6cd.cedula_antigua_anverso_v6vulnerability-cwe-patch-2data_ClassificationModel
Dataset Card for "data_ClassificationModel"---
dataset_info:
features:
- name: brands
dtype: string
- name: categories
dtype: string
- name: code
dtype: string
- name: languages_tags
dtype: string
- name: last_modified_t
dtype: int64
- name: product_name_de
dtype: string
- name: quantity
dtype: string
- name: index_level_0
dtype: int64
splits:
- name: train
num_bytes: 231023
num_examples: 673
download_size:… See the full description on the dataset page: https://huggingface.co/datasets/CedRuiz/data_ClassificationModel.task1663_cedr_ru_incorrect_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1663_cedr_ru_incorrect_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1663_cedr_ru_incorrect_classification.
