text-mining
text-mining-ce-dataset
Vietnamese Legal Cross-Encoder Dataset
Training data for a cross-encoder reranker on Vietnamese legal documents.
Source
Built from YuITC/Vietnamese-Legal-Documents.
Schema
Column
Type
Description
qid
int64
Query ID
cid
int64
Document (context) ID
query
string
Legal question
document
string
Candidate document
label
int64
1 = positive, 0 = negative
split
string
train or test
negative_type
string
random, same_topic_wrong_article… See the full description on the dataset page: https://huggingface.co/datasets/juzharii/text-mining-ce-dataset.TextMining-Phase-1-Index
phase1-indexes
Retrieval indexes built by the NewsQA RAG Phase 1 pipeline, packaged so a later
session can attach them instead of spending GPU time rebuilding.
Generated 2026-09-06T11:52:10+00:00 from /kaggle/input/notebooks/meowluvmatcha/newsqa-rag-phase-1-retrieval-tournament-kaggl/newsqa_phase1/indexes/round1.
Contents
round1/ — dense: BAAI/bge-large-en-v1.5, dense: BAAI/bge-small-en-v1.5, dense: all-MiniLM-L6-v2, dense: intfloat/e5-base-v2, sparse:… See the full description on the dataset page: https://huggingface.co/datasets/MatchaMacchiato/TextMining-Phase-1-Index.text-mining-ce-dataset-v2text-mining-rag-resultsRole-Mining-JSON_Inst-Texttext-mining-ce-dataset-v3
