CoolFace
3 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.tabularquestion-answeringn<1K0 likes91 downloads2y agoHugging Face02Loctran123 /vietnamese-evidence-corpus-chunked-e5-v3 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics Chunked with multilingual-E5 token budget Prefix-aware chunking using `passage: {title} ` Sentence-aware overlap to preserve local context Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.tabulartext-retrieval10K<n<100K0 likes39 downloads1mo agoHugging Face03Loctran123 /vietnamese-evidence-corpus-chunked-e5-v2 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.tabulartext-retrieval10K<n<100K0 likes20 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.