CoolFace
Datasetpublic

NoeFlandre/benchmark-llms-landuse-relevance

Land-use relevance benchmark v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels. Code Task and prompt Does a sentence describe a place's land or environment in ways visible to satellites? English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS GGUF quant through llama.cpp (same prompt, template, greedy decoding and budget).… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/benchmark-llms-landuse-relevance.

sourceHugging Facemitupdated 3d agoView on Hugging Face
0likes1.7kdownloads
Dataset Card

Land-use relevance benchmark

v3-multilingual · 85 languages x 300 items/language · 25,500 items · binary yes/no labels.

Code

Task and prompt

Does a sentence describe a place's land or environment in ways visible to satellites?

English prompt · greedy decoding · seed 0 · max_new_tokens=4096 · bfloat16 · batch varies by model. unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS runs the UD-IQ2_XXS GGUF quant through llama.cpp (same prompt, template, greedy decoding and budget).

Prompt text

Replace {} with the target sentence.

text
Classify whether the TARGET SENTENCE contains information about the target place that could help characterize its land use, land cover, or geographic environment from remote sensing, either directly or through observable proxies.

Return exactly one token: yes or no.

Answer yes for information about vegetation, agriculture, forests, water, soil or surface, terrain, buildings, settlements, infrastructure, transport networks, mining, managed land, or other human or natural features with a spatial or remotely detectable signature.

Answer no for information only about history, administration, people, events, demographics, economy, navigation, or activities with no meaningful land-use, land-cover, or remotely detectable implication.

Output only the lowercase token yes or no.

TARGET SENTENCE: {}

Aggregate scores

Per-model macro averages across languages. Full metrics: `aggregates.csv`. Bold = best; underline = second best in each metric column.

model_idlanguage_countaccuracy_macrobalanced_accuracy_macrof1_macroprecision_macrorecall_macromatthews_corrcoef_macro
Qwen/Qwen3.5-9B850.7851<u>0.7775</u>0.81510.75290.89190.5786
LiquidAI/LFM2.5-2.6B850.76020.7536<u>0.8029</u>0.72530.90010.5352
Qwen/Qwen3-4B-Instruct-2507850.76220.7550.79050.73560.8640.5344
Qwen/Qwen3-4B850.74720.7450.76340.75640.77840.4965
google/gemma-4-E4B-it85<u>0.7728</u>0.780.75090.8730.6719<u>0.5726</u>
Qwen/Qwen3-8B850.75530.76060.74420.8310.68180.5283
Qwen/Qwen3.5-0.8B850.61160.58820.72170.5883<u>0.939</u>0.2393
LiquidAI/LFM2.5-8B-A1B850.70470.70560.71790.73180.70960.4128
LiquidAI/LFM2.5-350M850.53330.50.69570.53331.00.0
Qwen/Qwen3-0.6B850.53330.50.69570.53331.00.0
Qwen/Qwen3.5-4B850.68330.69990.59360.90930.45150.4522
mistralai/Ministral-3-8B-Instruct-2512-BF16850.66880.68690.56250.91640.41560.4344
google/gemma-4-E2B-it850.63890.65770.52090.87460.37680.3732
LiquidAI/LFM2.5-1.2B-Instruct850.60380.61580.5160.71010.43510.2492
microsoft/Phi-4-mini-instruct850.60260.62210.45980.82950.32970.3005
HuggingFaceTB/SmolLM3-3B850.60710.62770.45510.86190.31870.3207
tiiuae/Falcon3-3B-Instruct850.59090.60990.44320.79350.32610.2673
ibm-granite/granite-3.3-2b-instruct850.56570.58460.40610.72130.30130.2016
unsloth/Qwen3.8-27B-GGUF@UD-IQ2_XXS850.58630.61130.3664<u>0.9521</u>0.23630.321
allenai/Olmo-3-7B-Instruct850.56880.59290.34690.85650.23150.2582
Qwen/Qwen3.5-2B850.56130.58640.33060.87390.2090.2569
mistralai/Ministral-3-3B-Instruct-2512-BF16850.55610.58150.31730.88110.2010.2507
utter-project/EuroLLM-9B-Instruct-2512850.52760.55640.21250.92130.1240.2129
tiiuae/Falcon3-7B-Instruct850.51390.54380.16570.91320.0950.1806
ibm-granite/granite-4.1-3b850.51110.54150.15470.97290.08630.1864
allenai/OLMo-2-1124-7B-Instruct850.50350.53410.13240.9290.07540.1543
Qwen/Qwen3-1.7B850.48930.52090.08430.76280.04680.1093
tiiuae/Falcon-H1-3B-Instruct850.48780.52020.080.8650.04280.1179
swiss-ai/Apertus-8B-Instruct-2509850.47610.50890.03460.88090.01780.0813
tiiuae/Falcon3-1B-Instruct850.47050.50350.01910.4180.00990.0263

Scoring models

Scores are normalized to [0, 1]. Best thresholds are selected on this benchmark (an upper bound); ROC-AUC needs no threshold. Full sweep: `threshold_sweep.csv`.

Scoring setup

modelhandlingrelevance score / decision rulesequence length (tokens)dtype / batch / seedrevision
Alibaba-NLP/gte-multilingual-reranker-basesequence classifier: prompt + sentence pairsigmoid relevance logit; yes if score ≥ 0.58192bfloat16; batch 16; seed 08215cf04918ba6f7b6a62bb44238ce2953d8831c
BalaRajesh1/mmbert-small-nliNLI zero-shot pipeline: sentence premise + hypothesisentailment probability; yes if score ≥ 0.5512bfloat16; batch 16; seed 02e7a7a1b86760ec00c5596f86e5b8bcb9ff8d9fb
LiquidAI/LFM2.5-2.6B@logprobcausal LM, no decoding: LLM prompt + chat turn, empty think blockfirst-token P(yes) vs P(no); argmax over the first-token yes/no log-probabilities8192bfloat16; batch 16; seed 0654f9463ce32b05d0429d76fe1f580b27d4c1ac0
MoritzLaurer/bge-m3-zeroshot-v2.0NLI zero-shot pipeline: sentence premise + hypothesisentailment probability; yes if score ≥ 0.5512bfloat16; batch 16; seed 09abf1c8aaeb82a2447809c20753ed0b106b76652
MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli-2mil7NLI zero-shot pipeline: sentence premise + hypothesisentailment probability; yes if score ≥ 0.5512bfloat16; batch 16; seed 0b5113eb38ab63efdd7f280f8c144ea8b13f978ce
Qwen/Qwen3-Reranker-0.6Bcausal-LM reranker: manual yes/no reranker turnyes/no next-token probability; argmax over the native yes/no scoresmodel-definedbfloat16; batch 16; seed 0e61197ed45024b0ed8a2d74b80b4d909f1255473
Qwen/Qwen3-Reranker-4Bcausal-LM reranker: manual yes/no reranker turnyes/no next-token probability; argmax over the native yes/no scoresmodel-definedbfloat16; batch 16; seed 022e683669bc0f0bd69640a1354a6d0aebcfeede5
convaiinnovations/laya-multilingualtyped decision model: JSON state + 4 noul questions/callLaya noul yes probability; yes if score ≥ 0.51024varies across runs (dtype is recorded per run); batch 16; seed 0b4a904d1a2a54c822b829e24291d4b8f280fe43e
fastino/gliner2.5-multi-v1GLiNER2 classify_text: sentence + hypothesis labellabel confidence; yes if score ≥ 0.5model-definedfloat32; batch 16; seed 0a221b77a8baf4a613b8f8652661d41fa10a5641e
knowledgator/gliclass-multilang-miniGLiClass zero-shot: sentence + hypothesis labellabel probability; yes if score ≥ 0.5512bfloat16; batch 16; seed 00bd888b6c3ef9fca5f0a9d407bddfbbc7623486b
mixedbread-ai/mxbai-rerank-base-v2causal-LM reranker: official query/document turnsigmoid(1-logit - 0-logit - 4.5); yes if score ≥ 0.58192bfloat16; batch 16; seed 03ea9d4dffa7d12a4f366be8e275c349de9fc9865

Scoring prompts

Alibaba-NLP/gte-multilingual-reranker-base, Qwen/Qwen3-Reranker-0.6B, Qwen/Qwen3-Reranker-4B, convaiinnovations/laya-multilingual, mixedbread-ai/mxbai-rerank-base-v2:

text
<Instruct>: Judge whether the Document carries information about its target place that could help characterize land use, land cover, or the geographic environment from remote sensing, either directly or through observable proxies. Answer yes for vegetation, agriculture, forests, water, soil or surface, terrain, buildings, settlements, infrastructure, transport networks, mining, managed land, or other human or natural features with a spatial or remotely detectable signature. Answer no for information only about history, administration, people, events, demographics, economy, navigation, or activities with no meaningful land-use, land-cover, or remotely detectable implication.
<Query>: Does this sentence carry land-use, land-cover, or geographic-environment signal observable from remote sensing?
<Document>: {}

BalaRajesh1/mmbert-small-nli, MoritzLaurer/bge-m3-zeroshot-v2.0, MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli-2mil7, fastino/gliner2.5-multi-v1, knowledgator/gliclass-multilang-mini:

text
Zero-shot classification input for NLI and label-matching models. The premise is the TARGET SENTENCE; the hypothesis / positive label is the line after HYPOTHESIS, derived from the LLM prompt.

TARGET SENTENCE: {}

HYPOTHESIS: This sentence contains information useful for inferring the land use or land cover of the associated geographic place.

LiquidAI/LFM2.5-2.6B@logprob uses the task prompt above.

Best thresholded scoring metrics

modellanguagesMCC @ thresholdF1 @ thresholdbalanced accuracy @ thresholdprecision @ thresholdrecall @ thresholdROC-AUCitems/speak VRAM (GiB)
LiquidAI/LFM2.5-2.6B@logprob850.3942 @ 0.90.7307 @ 0.80.6884 @ 0.90.7109 @ 0.91 @ 00.7773112.486.76
Qwen/Qwen3-Reranker-4B85<u>0.2897 @ 0.03</u><u>0.723 @ 0.003</u><u>0.6435 @ 0.03</u><u>0.8288 @ 0.3</u>1 @ 0<u>0.7051</u>n/an/a
MoritzLaurer/bge-m3-zeroshot-v2.0850.2573 @ 0.10.7041 @ 0.030.6226 @ 0.10.8344 @ 0.61 @ 00.6651112.571.25
knowledgator/gliclass-multilang-mini850.1987 @ 0.20.7037 @ 0.030.5976 @ 0.20.6383 @ 0.31 @ 00.6258258.37n/a
Qwen/Qwen3-Reranker-0.6B850.187 @ 0.00030.7026 @ 1e-050.593 @ 0.00030.7863 @ 0.031 @ 00.6353n/an/a
mixedbread-ai/mxbai-rerank-base-v2850.1416 @ 0.0030.7007 @ 0.0010.5608 @ 0.010.6345 @ 0.011 @ 00.615314.347.14
convaiinnovations/laya-multilingual850.0383 @ 0.40.6959 @ 0.010.5056 @ 0.40.5361 @ 0.41 @ 00.4871156.371.55
Alibaba-NLP/gte-multilingual-reranker-base850.0593 @ 0.50.6957 @ 00.5202 @ 0.60.5665 @ 0.61 @ 00.540978.641.29
BalaRajesh1/mmbert-small-nli850 @ 00.6957 @ 00.5 @ 00.5333 @ 01 @ 00.4992373.450.42
MoritzLaurer/mDeBERTa-v3-base-xnli-multilingual-nli-2mil7850.005 @ 0.90.6957 @ 00.5009 @ 0.90.5338 @ 0.91 @ 00.5284<u>260.91</u><u>0.74</u>
fastino/gliner2.5-multi-v1850.1305 @ 0.90.6957 @ 00.5607 @ 0.90.5833 @ 0.91 @ 00.607561.971.18

LFM2.5-2.6B: log-probabilities vs generation

Same model, prompt and languages. Generation can reason inside LFM's open <think> block before parsing its final yes/no; log-probs close the block empty and compare immediate next-token P(yes) vs. P(no), so the decision contexts differ.

methodF1MCCunparsedROC-AUCGPU hoursms/item
generation + parsing0.80290.53520.0042n/a40.115663.1
yes/no log-probs0.69650.01910.00.77730.069.0
same GPU (NVIDIA A100-SXM4-40GB; en, fr, zh): generation vs log-probs2462 s vs 7.1 s (349x)