CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aysinghal /cochange-code-retrieval Co-Change Code Retrieval A repository-level code retrieval dataset built from what real developers actually do together. Positives are labeled by co-change — files that engineers repeatedly modified in the same commits — not by imports, folder co-location, or any other structural proxy. Every training row ships with 128 hard + 128 easy negatives, where hard negatives are mined in the embedding space of a strong open model (Qwen3-Embedding-4B) and filtered through an eight-signal… See the full description on the dataset page: https://huggingface.co/datasets/aysinghal/cochange-code-retrieval.textfeature-extraction100K<n<1M0 likes493 downloads3mo agoHugging Face02aysinghal /code-retrieval-training-datasettext100K<n<1M1 likes387 downloads6mo agoHugging Face03benjamintli /code-retrieval-combined-v2text100K<n<1M0 likes68 downloads6mo agoHugging Face04AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.texttext-retrieval100K<n<1M0 likes29 downloads7mo agoHugging Face05AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-python RLVR-Env-Retrieval-Source-code-search-net-python RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.texttext-retrieval100K<n<1M0 likes24 downloads7mo agoHugging Face06benjamintli /code-retrieval-combined-v2-mined-negativestext1M<n<10M0 likes12 downloads6mo agoHugging Face07benjamintli /code-retrieval-hard-negatives-llm-verified-mergedtext100K<n<1M0 likes10 downloads6mo agoHugging Face08future7 /code_r1_with_retrievaltextn<1K0 likes7 downloads1y agoHugging Face09benjamintli /code-retrieval-combinedtext100K<n<1M0 likes7 downloads6mo agoHugging Face10benjamintli /code-retrieval-combined-v2-llm-negativestext1M<n<10M0 likes5 downloads6mo agoHugging Face11ahans1 /code_retrievalgatedtextn<1K0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.