code_retrieval
ide-code-retrieval-qwen3-0.6b-GGUFide-code-retrieval-qwen3-0.6bide-code-retrieval-gpt2-large-llm2vecide-code-retrieval-qwen3-0.6b-cochange-mask-hardmodernbert-python-code-retrievalide-code-retrieval-qwen3-0.6b-inbatchide-code-retrieval-qwen3-0.6b-ebs128SDWLLM-code_retrieval-MINDER-CodeT5-The_Vault_test_python
cochange-code-retrieval
Co-Change Code Retrieval
A repository-level code retrieval dataset built from what real developers actually do together. Positives are labeled by co-change — files that engineers repeatedly modified in the same commits — not by imports, folder co-location, or any other structural proxy. Every training row ships with 128 hard + 128 easy negatives, where hard negatives are mined in the embedding space of a strong open model (Qwen3-Embedding-4B) and filtered through an eight-signal… See the full description on the dataset page: https://huggingface.co/datasets/aysinghal/cochange-code-retrieval.code-retrieval-training-datasetcode-retrieval-combined-v2RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.code-retrieval-stackoverflow-small
