datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tab2-metricassmall-GPT-wiki-intro-features
Small-GPT-wiki-intro-features dataset
This dataset is based on aadityaubhat/GPT-wiki-intro.
It contains 100k randomly selected texts (50k from Wikipedia and 50k generated by ChatGPT).
For each text, various complexity measures were calculated, including e.g. readibility, lexical richness etc.
It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts.
Dataset structure
Features were calculated using… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/small-GPT-wiki-intro-features.tabela-taco
Dataset: TACO - Tabela Brasileira de Composição de Alimentos (Brazilian Food Composition Table)
The TACO table is a reference nutritional table for foods consumed in Brazil. The information contained in this dataset was taken from the excel file made available by NEPA - Center for Studies and Research in Food at UNICAMP, through the link:
https://nepa.unicamp.br/publicacoes/tabela-taco-excel/
License
The original TACO data remain subject to the terms established… See the full description on the dataset page: https://huggingface.co/datasets/julianamarques/tabela-taco.JuliaHealthDatasetsphishing-llm-bias-audit
LLM Phishing-Vulnerability Bias Audit Dataset
A multi-provider empirical dataset capturing how 14 open-source LLM configurations (across 5 inference providers) select which of three generated personas is "most vulnerable to phishing." 855 persona records / 285 forced-choice workflows.
Important. This dataset is about LLM behaviour under controlled prompts, not about real-world phishing susceptibility of any demographic group. Selecting a persona as "vulnerable" is the LLM's choice;… See the full description on the dataset page: https://huggingface.co/datasets/Julia569922/phishing-llm-bias-audit.GPT-wiki-intro-features
Small-GPT-wiki-intro-features dataset
This dataset is based on aadityaubhat/GPT-wiki-intro.
It contains 150k short texts from Wikipedia (label 0) and corresponding texts generated by ChatGPT (label 1) (together 300k texts).
For each text, various complexity measures were calculated, including e.g. readability, lexical diversity etc.
It can be used for text classification or analysis of linguistic features of human-generated and ChatGPT-generated texts.
For a smaller version… See the full description on the dataset page: https://huggingface.co/datasets/julia-lukasiewicz-pater/GPT-wiki-intro-features.yahoo-answersvk-media-tsrgetopenorca-10k-subset-llama-2mushroom_1.csvIMDb_movie_reviews
Dataset Card for IMDb Movie Reviews
Dataset Summary
This is a custom train/test/validation split of the IMDb Large Movie Review Dataset available from http://ai.stanford.edu/~amaas/data/sentiment/.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
IMDb_movie_reviews
An example of 'train':
{
"text": "Beautifully photographed and ably acted, generally, but the… See the full description on the dataset page: https://huggingface.co/datasets/julianuribe03/IMDb_movie_reviews.Rosberta_100diss_512chunksizeRosberta_alldiss_512chunksize
