retriever
ko-law-retriever-artifacts-20260622
Korean Legal Retriever Artifacts 2026-06-22
This public dataset repository stores large artifacts for ko-law-retriever
that are too large for GitHub's 100 MB file limit.
Contents
ko_legal_source_rag/: retriever SFT JSONL mixes and metadata.
ko_legal_retriever/ko_legal_fts_full.sqlite: local SQLite FTS retrieval index.
evals/: source-rag evaluation outputs.
reports/: analysis JSON/Markdown reports and evidence-pack outputs.
Scope
These files are… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ko-law-retriever-artifacts-20260622.azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.daily-paper-2026-07-09-retriever-vs-decomposition-skill-routing
Retriever Bottleneck vs. Decomposition
TL;DR — In a real 1,898-skill bilingual agent harness, the ceiling-gap diagnostic
shows that neither decomposition nor a better retriever raises routing coverage: the
binding constraint is skill-corpus redundancy. Decomposition still earns its place on
execution ordering.
ThakiCloud AI Research · 2026-07-09 (v2, revised) · 📝 Tech blog (KO)
Problem
Operators of large agent harnesses that route natural-language requests to… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-09-retriever-vs-decomposition-skill-routing.corpus.latest.vod-retriever-medical-v1.1cole-arxiv-cc-e5-retriever
CoLe arXiv CC E5 Retriever
A non-Wikipedia retrieval corpus for the CoLe/R2–R4a experiments. It contains
280,738 arXiv documents selected from the Common Pile
filtered arXiv collection, plus an E5-base-v2 dense index.
Data provenance and license handling
Source: common-pile/arxiv_papers_filtered, revision 033cf7f.
The source records are converted arXiv papers with per-document license metadata.
This release keeps records whose metadata is CC BY, CC0, or Public… See the full description on the dataset page: https://huggingface.co/datasets/KyrieX/cole-arxiv-cc-e5-retriever.retriever-princeton-nlp-CharXiv-clean
Description
princeton-nlp/CharXiv dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@article{wang2024charxiv,
title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs},
author={Wang, Zirui and Xia, Mengzhou and He, Luxi and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-princeton-nlp-CharXiv-clean.
