datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pmp-se-test-dataset
Dataset Card for Dataset Name
PMP-Synth-AllPairs-Annotatedpm-price-historiesprocessed-old-with-embeddings
Dataset Card for "processed-old-with-embeddings"
Dataset Summary
Chunks of about 256 words split by whitespace and their embeddings computed with the pretrained spacy model ["de_dep_news_trf"] (https://github.com/explosion/spacy-models/releases/tag/de_dep_news_trf-3.6.1).
The splits are created with respect to sentence boundaries parsed with the same model, sentences are concatenated if the result does not exceed max_words = 256, therefore the chunk length varies.… See the full description on the dataset page: https://huggingface.co/datasets/pmpc/processed-old-with-embeddings.
