datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cdncdnpdf-presentations-part1
Dataset Card for cdnpdf Educational Materials (Part 1)
Dataset Summary
This dataset contains metadata and original files for 101,022 educational presentations from the cdnpdf.com platform, which provides free access to books, documents, magazines and presentations. This collection focuses exclusively on presentations and includes archives with IDs from 00 to 24. The dataset includes information such as presentation titles, descriptions, URLs, download URLs, and file… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cdnpdf-presentations-part1.CDNA-CoDET-M4
CDNA-CoDET-M4: Code Authorship Attribution via Code Property Graphs (Enhanced)
Dataset Summary
Note: This is an enhanced version of the original CoDET-M4 dataset by DaniilOr, extended with Code Property Graph (CPG) representations by the CodeDNA team at Singapore Management University.
We built this dataset to tackle LLM code authorship attribution—figuring out exactly which AI model wrote a specific piece of code. While most approaches just analyze the raw source code… See the full description on the dataset page: https://huggingface.co/datasets/mohameddhameem/CDNA-CoDET-M4.cdnpdf-presentations-part2
Dataset Card for cdnpdf Educational Materials (Part 2)
Dataset Summary
This dataset contains metadata and original files for 100,978 educational presentations from the cdnpdf.com platform, which provides free access to books, documents, magazines and presentations. This collection focuses exclusively on presentations and includes archives with IDs from 25 to 49. The dataset includes information such as presentation titles, descriptions, URLs, download URLs, and file… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/cdnpdf-presentations-part2.cblue-cdnCDNA
Overview
Authors construct a Chinese LLM safety evaluation by translating and localizing the "Do-not-answer" dataset and expand it with region-specific questions and align it with country-specific AI generation regulations,
Authors then extend the resulting 1,014 questions from two prespectives:
False Negative(FN) questions: risky questions posed in an
evasive way, aimed at evaluating an LLM’s sensitivity to perceiving risks, aimed at evaluating an LLM’s sensitivity to perceiving… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CDNA.human_cdnadna cdna hamburger
Dataset Card for cdna_test_dset
Dataset Summary
This dataset is very nice
Supported Tasks and Leaderboards
[Needs More Information]
Languages
[Needs More Information]
Dataset Structure
Data Instances
this is how the data could look
{
'sequence':'ACTGGTTC',
}
Data Fields
[Needs More Information]
Data Splits
no splits yet
Dataset Creation
Curation Rationale
[Needs… See the full description on the dataset page: https://huggingface.co/datasets/Vlasta/human_cdna.CDN_es_encdnc_lawcdnc_law_test
Dataset Card for "cdnc_law_test"
More Information needed
cdnc_law_eval
Dataset Card for "cdnc_law_eval"
More Information needed
daft-sm89-cdna3-functionscdnmcz-cdndaft-sm90-cdna3-functions
