datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cnn-based-drowsiness-detection-data
CNN-Based Drowsiness Detection - Dataset
Preprocessed, auto-labeled face-crop images used to train the model in
notgoodkeeper/cnn-based-drowsiness-detection.
Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection
Collection
Frames were captured from a webcam, then run through:
Haar Cascade face detection -> crop + pad + resize to 412x412
MediaPipe Selfie Segmentation -> background replaced with white
CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.NO-CNN-DailyMail
Dataset Card
Dataset Summary
NO-CNN-DailyMail is a Norwegian news summarization dataset partially machine translated from English version of CNN Dailymail Dataset. The summaries were written by journalists at CNN and the DailyMail. The dataset can be used for Machine reading comprehension and abstractive summarization tasks.
Data Instances
For each instance, there is an article string and a positive_sample string representing news article and abstractive… See the full description on the dataset page: https://huggingface.co/datasets/NorGLM/NO-CNN-DailyMail.cn-news-impact-scores
Chinese News Impact Scores 2024-2025
This dataset pairs complete Chinese financial-news collections for 2024 and 2025 with event-level market, industry/board, and stock impact scores. Data is stored in monthly Parquet shards.
Dataset Structure
raw_news: every collected news occurrence, including full text, a unique occurrence_id, and a stable news_id.
impact_scores: one row per event-target pair with routing metadata and 16 impact dimensions.
event_clusters:… See the full description on the dataset page: https://huggingface.co/datasets/GlazJ/cn-news-impact-scores.CNN_NEWS_DATASETUDR_CNNDailyMail
Dataset Card for "UDR_CNNDailyMail"
More Information needed
cnn_daily_mail_traincnndm_10k_semantic_rouge_labelsbart_cnn_lyric_summariessummarize_from_feedback_oai_preprocessing_1706381144_cnndm_relabel_pythia6.9b_emojijoint-rams-and-cnn-v1cnn_dm_paraphrase_10k
Dataset Card for "cnn_dm_paraphrase_10k"
More Information needed
cnn_dm_paraphrase_full
Dataset Card for "cnn_dm_paraphrase_full"
More Information needed
bitcoin_oversample_cnn_extractedshivam9980__mistral-7b-news-cnn-merged-details
Dataset Card for Evaluation run of shivam9980/mistral-7b-news-cnn-merged
Dataset automatically created during the evaluation run of model shivam9980/mistral-7b-news-cnn-merged
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/shivam9980__mistral-7b-news-cnn-merged-details.cnn_dm_paraphrase_small
Dataset Card for "cnn_dm_paraphrase"
More Information needed
shivam9980__mistral-7b-news-cnn-mergedsummarize_from_feedback_oai_preprocessing_1706381144_cnndm_relabel_pythia6.9bCNNGRAPH-datasetcnn-nppe-datasetcnndm_llama2_7b_chat_summary
Dataset Card for "cnndm_llama2_7b_chat_summary"
More Information needed
mnist_cnncnn_dm_paraphrase_unique_50k
Dataset Card for "cnn_dm_paraphrase_unique_50k"
More Information needed
summarize_from_feedback_oai_preprocessing_1706381144_cnndm_relabel_pythia6.9b_all_prefixescnndailymail_extractivecnn-dailymail-processedcnn-dm-human-evalfrom https://github.com/whl97/LS-Score
Scratch_CNN-logs-datasetPretrained_CNN-logs-datasetEFFICENTNET_CNN-logs-datasetResNet152_CNN-logs-dataset
