anhaidgroup/polaris-ecir-v1
ECIR v1 ECIR is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval Feedback, alongside aw, arctic, lter, wikitables, and wtr. It holds 2,100 tables published on the US government open data portal and 12 keyword queries over them. For each query–table pair, a person decided how well that table answers that query and gave it a score; those scores are the relevance judgments, and they live in qrels.csv. Given a query, a system ranks the 2,100… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-ecir-v1.
ECIR v1
ECIR is one of six datasets in [Polaris: Learning to Generate Table Descriptions from Retrieval Feedback](https://arxiv.org/abs/2608.17171), alongside aw, arctic, lter, wikitables, and wtr.
It holds 2,100 tables published on the US government open data portal and 12 keyword queries over them. For each query–table pair, a person decided how well that table answers that query and gave it a score; those scores are the relevance judgments, and they live in qrels.csv. Given a query, a system ranks the 2,100 tables, and the judgments say which ones should have come back.
The tables have no names. What describes a table is its column names and the portal listing it was published under — the dataset title, the notes, and a short description.
Each Polaris dataset comes in two versions. v1, this one, has the table metadata, the queries, and the relevance judgments. v2 adds the tuples — the contents of each table — and is at polaris-ecir-v2.
Files
`queries.csv` — one row per query. There are 12 queries but only 6 topics: Chen et al. wrote two wordings for each topic and scored tables against the topic rather than the wording. So q1 and q2 carry identical rows in qrels.csv, as do q3/q4 and so on, and anything averaged over all 12 queries counts each topic twice.
query_id,query
q1,Wind speed in Kansas in years 2003-2004
q2,Kansas historical weather records`qrels.csv` — one row per query–table pair that a person scored. Below, two of the tables scored for q1.
query_id,table_id,relevance_score
q1,ECIR-1375,3.0
q1,ECIR-305,0.44175960347Scores are floats because crowdworkers rated each pair 0 to 3 and the platform weighted each worker by how reliable they had proven to be, so the result lands on 0.44175960347 rather than 0.33. Do not round it; anything above 0 is relevant.
Judging was done by pooling — for each query only a set of candidate tables was scored, not the whole corpus. So a 0 means someone looked and said no, while a missing pair means nobody looked. 1,676 of the 2,100 tables appear here; the other 424 do not.
`metadata.csv` — one row per table. Below is ECIR-1375, which scored 3.0 for q1.
table_id,column_names,table_context
ECIR-1375,"[""Read by"", ""DR Version 10"", ""10/14/2003 10:27"", ""Unnamed: 3""]","{""dataset_title"": ""Anemometer Data (Wind Speed, Direction) for Beloit, Kansas (2003 - 2004)"", ...}"table_context holds the portal's dataset_title, dataset_notes, and table_description.
Statistics
A table counts as gold if it scores above 0 for at least one query.
ECIR is by far the densest of the six — roughly a fifth of the corpus is relevant to any given query.
Download
The three files come to about 3.3 MB. You do not need a Hugging Face account to download them.
Option 1 — click the Files tab at the top of this page and save each file.
Option 2 — command line (recommended):
pip install huggingface_hub
hf download anhaidgroup/polaris-ecir-v1 --repo-type dataset --local-dir ecirPolaris has six datasets in total and ECIR is one of them. Each sits in its own repository, so to download all six quickly — the v1 repositories, without tuples:
for d in aw arctic lter ecir wikitables wtr; do
hf download anhaidgroup/polaris-$d-v1 --repo-type dataset --local-dir polaris_v1/$d
doneUsage
column_names and table_context are both arrays or objects, so they need parsing when you load the file. For example:
import ast
import pandas as pd
metadata = pd.read_csv("ecir/metadata.csv")
queries = pd.read_csv("ecir/queries.csv")
qrels = pd.read_csv("ecir/qrels.csv")
metadata["columns"] = metadata["column_names"].apply(ast.literal_eval)
metadata["context"] = metadata["table_context"].apply(ast.literal_eval)
relevant = qrels.loc[(qrels.query_id == "q1") & (qrels.relevance_score > 0), "table_id"].tolist()Note the > 0: scores are fractional, so an exact test like == 1 would match almost nothing.
How is this dataset created?
The tables and the original judgments come from a dataset-search benchmark built on the US government open data portal by Chen et al. (ECIR 2020). That benchmark defines 6 search tasks, each with 20 queries, and scores task–table pairs on a 0–3 scale using judgments from multiple annotators. Earlier work treats all 120 queries as independent.
The Polaris authors found two problems with it. The 20 queries within a task are highly similar and lack diversity, and some are unrelated to their task. And the labels have false positives: many task–table pairs scored 2 or above turn out not to be relevant on inspection, while the pairs scored 0 are generally correct.
They fixed both. Every pair originally scored 2 or above was manually re-labelled, and the 0s were left as they were. To cut the redundancy inside each task, 2 representative and maximally distinct queries were kept per task, leaving 12. Tables that are not CSV or cannot be opened were removed, leaving 2,100.
Citation
@misc{cai2026polaris,
title = {Polaris: Learning to Generate Table Descriptions from Retrieval Feedback},
author = {Cai, Ting and Phan, Tuan Minh and Doan, AnHai},
year = {2026},
eprint = {2608.17171},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
doi = {10.48550/arXiv.2608.17171},
url = {https://arxiv.org/abs/2608.17171}
}Please also cite the collection it is built on:
@inproceedings{chen2020ecir,
author = {Zhiyu Chen and Haiyan Jia and Jeff Heflin and Brian D. Davison},
title = {Leveraging Schema Labels to Enhance Dataset Search},
booktitle = {Proceedings of the European Conference on Information Retrieval ({ECIR})},
pages = {267--280},
year = {2020},
}License
Contact
Email minhrua@cs.wisc.edu, valid until May 2029. After that, email anhai@cs.wisc.edu.
