datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yt-transcriptionswikipedia-me5Cohere's Simple Wikipedia Embedded with Multilingual E5 Large
sst2-stats
Stats for stanfordnlp/sst2
Generated by dataset-stats.py over the train split.
Total rows in split: 67,349
Rows profiled: 5,000
Columns: 3
Column overview
column
type
kind
null %
highlights
idx
Value(int32)
numeric
0.0%
min=0.00 · p50=2,499.50 · max=4,999.00 · distinct=5000
sentence
Value(string)
string
0.0%
distinct=4,999 · len p50=39 · max=255
label
ClassLabel
class_label
0.0%
positive=2758 · negative=2242
Per-column detail… See the full description on the dataset page: https://huggingface.co/datasets/Tim-Pinecone/sst2-stats.
