MrAiran/nsfw-pt_br
Dataset Card for The Pile Dataset Summary The NSFW is a 230K diverse, filtred and cleaned text from adult websites, high-quality datasets combined together. Supported Tasks and Leaderboards [More Information Needed] Languages This dataset is in Portuguese Brazil (pt_BR) Dataset Structure Data Instances Data Fields all text (str): Text. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/MrAiran/nsfw-pt_br.
Dataset Card for The Pile
Table of Contents
- Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Summary
The NSFW is a 230K diverse, filtred and cleaned text from adult websites, high-quality datasets combined together.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
This dataset is in Portuguese Brazil (pt_BR)
Dataset Structure
Data Instances
Data Fields
all
text(str): Text.
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data
Initial Data Collection and Normalization
[More Information Needed]
Who are the source language producers?
[More Information Needed]
Annotations
Annotation process
[More Information Needed]
Who are the annotators?
[More Information Needed]
Personal and Sensitive Information
[More Information Needed]
Considerations for Using the Data
Social Impact of Dataset
[More Information Needed]
Discussion of Biases
[More Information Needed]
Other Known Limitations
[More Information Needed]
Additional Information
Dataset Curators
This dataset was primarily curated by Airan with several adult websites
Licensing Information
Please refer to the specific license depending on the subset you use:
- Free to use
