CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aysinghal /cochange-code-retrieval Co-Change Code Retrieval A repository-level code retrieval dataset built from what real developers actually do together. Positives are labeled by co-change — files that engineers repeatedly modified in the same commits — not by imports, folder co-location, or any other structural proxy. Every training row ships with 128 hard + 128 easy negatives, where hard negatives are mined in the embedding space of a strong open model (Qwen3-Embedding-4B) and filtered through an eight-signal… See the full description on the dataset page: https://huggingface.co/datasets/aysinghal/cochange-code-retrieval.textfeature-extraction100K<n<1M0 likes492 downloads2mo agoHugging Face02aysinghal /code-retrieval-training-datasettext100K<n<1M1 likes388 downloads6mo agoHugging Face03benjamintli /code-retrieval-combined-v2text100K<n<1M0 likes68 downloads6mo agoHugging Face04AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.texttext-retrieval100K<n<1M0 likes32 downloads7mo agoHugging Face05AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-python RLVR-Env-Retrieval-Source-code-search-net-python RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.texttext-retrieval100K<n<1M0 likes27 downloads7mo agoHugging Face06code-rag-bench /code-retrieval-stackoverflow-smalltext10K<n<100K0 likes17 downloads2y agoHugging Face07feedback-to-code /retrieval_bench_1textn<1K0 likes14 downloads3y agoHugging Face08benjamintli /code-retrieval-combined-v2-mined-negativestext1M<n<10M0 likes12 downloads6mo agoHugging Face09benjamintli /code-retrieval-hard-negatives-llm-verified-mergedtext100K<n<1M0 likes10 downloads6mo agoHugging Face10future7 /code_r1_with_retrievaltextn<1K0 likes7 downloads1y agoHugging Face11benjamintli /code-retrieval-combinedtext100K<n<1M0 likes7 downloads6mo agoHugging Face12zard1152 /retrieval_plugin_code0 likes5 downloads3y agoHugging Face13benjamintli /code-retrieval-combined-v2-llm-negativestext1M<n<10M0 likes5 downloads6mo agoHugging Face14jingyq1 /AR-RAG_Retrieval_Code0 likes2 downloads1y agoHugging Face15Tasfiya025 /Code-Text_Retrieval_and_Documentation_Generation0 likes1 downloads9mo agoHugging Face16ahans1 /code_retrievalgatedtextn<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.