user-simulation
user-simulation-results
User simulation leaderboard: results
One file per system per release, under <org>/<system>/results_<timestamp>.json.
results holds the score the leaderboard displays (Recall@10 per benchmark). record holds
everything that makes the row citable and is not shown in the grid: the person-clustered
bootstrap interval, n, NDCG@10, the popularity and stranger controls, the evaluation cell, and
the run that produced it (experiment id, commit, date, seed).
These files are generated… See the full description on the dataset page: https://huggingface.co/datasets/jean-technologies/user-simulation-results.user-simulation-requests
User simulation leaderboard: requests
One file per submitted system, under <org>/<system>_eval_request_*.json, written by the
leaderboard Space's submission form.
status is PENDING when submitted, FINISHED once Jean has rerun the system against the
frozen benchmarks. artifact records what Jean was given to rerun (open weights, a container or
script, an API endpoint, or nothing, in which case the entry is listed as reported and is never
ranked above a rerun row). verification… See the full description on the dataset page: https://huggingface.co/datasets/jean-technologies/user-simulation-requests.user-reply-modelling
User Reply Modelling
2 billion tokens of (recent post history : post that user replied to : user's reply to post) sets for deeply modelling user interactions. Collected between 2024-11-28 and 2024-12-04 on BlueSky.
