LAM
Datasets
All datasets matching “LAM”lambada_openai
Dataset Summary
This dataset is comprised of the LAMBADA test split as pre-processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian.
LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lambada_openai.vtok101-distr-attribution-baselines
vtok101 attribution baselines, with a hard negative beside every document
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-distr-lora-seeds:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-distr-attribution-baselines.lambada
Dataset Card for LAMBADA
Dataset Summary
The LAMBADA evaluates the capabilities of computational models
for text understanding by means of a word prediction task.
LAMBADA is a collection of narrative passages sharing the characteristic
that human subjects are able to guess their last word if
they are exposed to the whole passage, but not if they
only see the last sentence preceding the target word.
To succeed on LAMBADA, computational models cannot
simply rely on local… See the full description on the dataset page: https://huggingface.co/datasets/cimec/lambada.route-attribution-baselines
vtok101 attribution baselines
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-route-ab-l8-lora-scale:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking those documents come.
Every document in the pool… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/route-attribution-baselines.vtok101-attribution-baselines
vtok101 attribution baselines
Data-attribution scores over lamsheeper-data-attribution/Qwen3.5-4B-d0-vtok101-lora-seeds:
3 function counts x 7 document counts x 4 seeds,
scored by 12 methods.
Each training document defines one synthetic constant function, and each query
asks for one function's value. The ground truth for a query is the set of
documents describing its function, so a method is measured by how far up its
ranking those documents come.
Every document in the pool… See the full description on the dataset page: https://huggingface.co/datasets/lamsheeper-data-attribution/vtok101-attribution-baselines.lamthikieu1998
