CoolFace
Datasetpublic

NaverHustQA/TVPL

TVPL (thuvienphapluat.vn) structured_data_doc.parquet: preprocessed version from only needed doc from tvpl . Please read by Datasets library parent_nodes.parquet: parent nodes from [1] by chunking with SentenceSplitter, chunk_overlap=0, chunk_size=800, tokenizer="Viet-Mistral/Vistral-7B-Chat" child_nodes.parquet: child nodes from [2] by chunking with SentenceSplitter, chunk_overlap=30, chunk_size=190 and using Word Segmentation, vietnamese-bi-encoder Dedup… See the full description on the dataset page: https://huggingface.co/datasets/NaverHustQA/TVPL.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes54downloads
2 commits on main
9964ae32y ago

Upload folder using huggingface_hub

thanhnx12
110f2812y ago

initial commit

thanhnx12