Retriever
ko-law-retriever-artifacts-20260622
Korean Legal Retriever Artifacts 2026-06-22
This public dataset repository stores large artifacts for ko-law-retriever
that are too large for GitHub's 100 MB file limit.
Contents
ko_legal_source_rag/: retriever SFT JSONL mixes and metadata.
ko_legal_retriever/ko_legal_fts_full.sqlite: local SQLite FTS retrieval index.
evals/: source-rag evaluation outputs.
reports/: analysis JSON/Markdown reports and evidence-pack outputs.
Scope
These files are… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ko-law-retriever-artifacts-20260622.azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.daily-paper-2026-07-09-retriever-vs-decomposition-skill-routing
Retriever Bottleneck vs. Decomposition
TL;DR — In a real 1,898-skill bilingual agent harness, the ceiling-gap diagnostic
shows that neither decomposition nor a better retriever raises routing coverage: the
binding constraint is skill-corpus redundancy. Decomposition still earns its place on
execution ordering.
ThakiCloud AI Research · 2026-07-09 (v2, revised) · 📝 Tech blog (KO)
Problem
Operators of large agent harnesses that route natural-language requests to… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-09-retriever-vs-decomposition-skill-routing.corpus.latest.vod-retriever-medical-v1.1retriever-princeton-nlp-CharXiv-clean
Description
princeton-nlp/CharXiv dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@article{wang2024charxiv,
title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs},
author={Wang, Zirui and Xia, Mengzhou and He, Luxi and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-princeton-nlp-CharXiv-clean.Qwen2.5-7B-Instruct-vllm-retriever-20251202_093826
