CoolFace
20 results

mining

loicmagne /open-subtitles-bitext-miningtext1M<n<10M1 likes1.9k downloads2y agoHugging Faceloicmagne /open-subtitles-256s-bitext-miningtext100K<n<1M0 likes1.8k downloads2y agoHugging Facemteb /tatoeba-bitext-mining Tatoeba An MTEB dataset Massive Text Embedding Benchmark 1,000 English-aligned sentence pairs for each language based on the Tatoeba corpus Task category t2t Domains Written Reference https://github.com/facebookresearch/LASER/tree/main/data/tatoeba/v1 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["Tatoeba"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tatoeba-bitext-mining.texttranslation100K<n<1M9 likes1.6k downloads7mo agoHugging Facedataforge-labs /bitcoin-mining-pool-templates Bitcoin mining pool templates Timestamped Stratum job messages collected directly from Bitcoin mining pool endpoints. The data records changes in the work each endpoint sends to miners, including the previous block hash, coinbase data and clean-jobs flag. Contents Table Record bitcoin_mining_pool_jobs A job received from a pool endpoint, with its observation time, nTime, coinbase, merkle branch count and clean-jobs flag Using the data… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/bitcoin-mining-pool-templates.tabulartime-series-forecasting100K<n<1M0 likes1.6k downloads4h agoHugging Facemesolitica /instructions-pair-miningtext100K<n<1M2 likes1k downloads3y agoHugging FaceSaylorTwift /mteb-bitext-mining-aggregated MTEB BitextMining Aggregated Dataset (Full) This dataset aggregates ALL configs from 10 BitextMining datasets in the MTEB (Massive Text Embedding Benchmark) Multilingual v2 benchmark into a single, unified dataset for comprehensive bitext mining evaluation. Dataset Summary Total Examples: 448,229 sentence pairs Source Datasets (Configs): 10 MTEB BitextMining tasks Total Splits: 332 language pairs/configurations Languages: 300+ unique language codes across all datasets… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/mteb-bitext-mining-aggregated.textsentence-similarity100K<n<1M0 likes893 downloads6mo agoHugging Face