datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x
qwen35-action-only-10k — Terminal-Bench 2.1 (timeout multiplier 2x)
Terminal-Bench 2.1 evaluation protocol variant (timeout multiplier 2x) of violetxi/qwen35-4b-offline-echo-action-only-10k-tacc through the served
model ID qwen35-action-only-10k with Terminus-2.
Result
Evaluation trials: 445
Tasks / attempts: 89 × 5
Errored trials scored as zero: 219
Agent timeouts / context-length events / output-cap events:
217 / 0 /
0
Mean reward / Pass@1: 0.105618
Pass@5:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-4b-offline-echo-action-only-10k-tacc-timeout2x.Echo88-150M-Base
Echo88-150M-Base
Echo88-150M-Base is a small English decoder-only causal language model trained from scratch on the Echo88 pretraining dataset.
The goal of Echo88 is to create a compact base model inspired by the language, computing culture, printed media, Usenet discussion, and older book knowledge available up to the late 1980s.
This is a base model, not an instruction-tuned chatbot. It is trained for next-token prediction and should be fine-tuned before being used as a… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/Echo88-150M-Base.mlabonne__BigQwen2.5-Echo-47B-Instruct-details
Dataset Card for Evaluation run of mlabonne/BigQwen2.5-Echo-47B-Instruct
Dataset automatically created during the evaluation run of model mlabonne/BigQwen2.5-Echo-47B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/mlabonne__BigQwen2.5-Echo-47B-Instruct-details.Echo-datasets
Echo RAG Eval
Index
Use wiki18_100w as the database with the E5-base-v2 embedder.
The retrieval corpus is large and must be downloaded manually. Download wiki18_100w from ModelScope:
https://www.modelscope.cn/datasets/hhjinjiajie/FlashRAG_Dataset/tree/master/retrieval_corpus
we have prepared the download script, after it, put the corpus under tests/data/retrieval_corpus, including wiki18_100w.jsonl and e5_flat_inner.index
the corpus has prebuilt FAISS index with… See the full description on the dataset page: https://huggingface.co/datasets/Ecoka/Echo-datasets.EchoMist
Dataset Card for EchoMist
Introducing EchoMist, the first comprehensive benchmark to measure how LLMs may inadvertently Echo and amplify Misinformation hidden within seemingly innocuous user queries.
Dataset Description
Prior work has studied language models' capability to detect explicitly false statements. However, in real-world scenarios, circulating misinformation can often be referenced implicitly within user queries. When language models tacitly agree, they may… See the full description on the dataset page: https://huggingface.co/datasets/ruohao/EchoMist.liberty-echo-voice-assetsNorquinal__Echo-details
Dataset Card for Evaluation run of Norquinal/Echo
Dataset automatically created during the evaluation run of model Norquinal/Echo
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Norquinal__Echo-details.echo-ml-artifactscadenza-echoblast-deception-evalGated research artifact. Request access for details.
