CoolFace
Datasetpublic

bkai-foundation-models/NewsSapo

Vietnamese NewsSapo Dataset The Vietnamese NewsSapo dataset was constructed to train sentence/passage embeddings. Our dataset is structured in a "title-abstract-contents" format, where each news article is represented by a tuple of (title, abstract, content). The content is the main text body of the article and has been processed to remove images, videos, and other non-textual elements. The dataset contains 31,728,183 triples. To build this dataset, we followed a two-step process: Step 1:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/NewsSapo.

sourceHugging Faceupdated 3y agoView on Hugging Face
6likes417downloads
8 commits on main
e6b69703y ago

Update README.md

sangdv
ceb971f3y ago

Update README.md

sangdv
31d1d103y ago

Update README.md

sangdv
de1ba3f3y ago

Update README.md

sangdv
ac1dda13y ago

Update README.md

iambestfeed
c5c343a3y ago

Create README.md

iambestfeed
9de59893y ago

Upload folder using huggingface_hub

iambestfeed
f0d4dfc3y ago

initial commit

iambestfeed