locale
Datasets
All datasets matching “locale”locale-benchmark-sra500
Locale Embedding Benchmark : SRA 500
What
This benchmark contains embeddings of raw genomic read sequences produced by the LOCALE DNA transformer model to test the use of vector search over large sequence repositories like the NIH Sequence Read Archive. The benchmark contains the embeddings of 163,578,486 sequence embeddings coming from 500 SRA Accessions. All vectors are Float32, D=768, roughly 500GB of total data.
Given a read, we want to find accessions… See the full description on the dataset page: https://huggingface.co/datasets/rsynk/locale-benchmark-sra500.local_eval
Local Affine duel (remote vLLM, Affine scoring)
This machine does not serve models. Affine's evalsrv.dueling.run_duel scores
three remote OpenAI-compatible endpoints. Model ids come from each server's
GET /v1/models — you only pass URLs. Completions still send that served id.
Chat tokenizers are loaded from tokenizer/{teacher,king,challenge}
(tokenizer.json + chat_template.jinja), not Hugging Face.
teacher http://HOST:10000
king http://HOST:10001
challenger… See the full description on the dataset page: https://huggingface.co/datasets/0xbidkslj1/local_eval.local-embeddings-2022
Local Embeddings Dataset
Multi-temporal satellite imagery dataset for phenology embedding training.
Dataset Description
This dataset contains multi-spectral satellite tiles across 6 months (April-September 2022) with 16 bands per tile.
Dataset Structure
local_embeddings/
├── alphaearth_embeddings_tiles_202204/ (263 tiles)
├── alphaearth_embeddings_tiles_202205/ (263 tiles)
├── alphaearth_embeddings_tiles_202206/ (263 tiles)
├──… See the full description on the dataset page: https://huggingface.co/datasets/gabrielireland/local-embeddings-2022.2A2B2C-Dataset-45D-LocalEElocale-benchmark-sra50
Locale Embedding Benchmark : SRA 50
What
This benchmark contains embeddings of raw genomic read sequences produced by the LOCALE DNA transformer model to test the use of vector search over large sequence repositories like the NIH Sequence Read Archive. The benchmark contains the embeddings of 9,688,220 sequence embeddings coming from 50 SRA Accessions. All vectors are Float32, D=768, roughly 30GB of total data.
Given a read, we want to find accessions containing… See the full description on the dataset page: https://huggingface.co/datasets/rsynk/locale-benchmark-sra50.PersonaSignal-LeakageCheck-Locale-And-Time-Zone-gpt-4o-mini
