datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ValidateSplitannual_reportsThe dataset is a comprehensive collection of financial documents from corporate annual reports.
Reference: https://www.annualreports.com/
attempt6_hm10MM-Mind2Web-tilde_test_snapshot_20dist
MultiModal-Mind2Web~ (MM-Mind2Web~)
rabbit inc.
[Leaderboard & Blogpost to be released]
Configuration: test split, snapshot with seed 42, 20 distractors
Multimodal-Mind2Web is a dataset proposed by Boyuan et al.. It's designed for the development and evaluation of generalist web agents and includes various action trajectories of humans on real websites.
We've simplified the raw dump from both Multimodal-Mind2Web and Mind2Web into sequences of observation-action pairs. We've… See the full description on the dataset page: https://huggingface.co/datasets/rabbit-hmi/MM-Mind2Web-tilde_test_snapshot_20dist.attempt7_hm5btc_20260818T1100.v2.1s.parquet
BTC Up/Down 5m — one-second bars, market and underlying on one clock
Polymarket runs a Bitcoin Up or Down market every five minutes, resolved off a Chainlink
TWAP of BTC/USD. This is one hour of that market on a one-second grid — and, on the same
grid, the spot and perpetual books from the four exchanges the price comes from, plus the
oracle stream the market resolves against.
Every grid is complete and gapless, and the winning outcome is already joined onto the… See the full description on the dataset page: https://huggingface.co/datasets/RabbitGo777/btc_20260818T1100.v2.1s.parquet.emgena_python_distributed_celery_rabbitmq_triage_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_python_distributed_celery_rabbitmq_triage_teaser.llama3datarabbit_b4bqaThis is the drug-matching dataset between generic and brand keywords for the RABBIT leaderboard 🐰.
And here is the paper: arxiv
@misc{gallifant2024language,
title={Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks},
author={Jack Gallifant and Shan Chen and Pedro Moreira and Nikolaj Munch and Mingye Gao and Jackson Pond and Leo Anthony Celi and Hugo Aerts and Thomas Hartvigsen and Danielle Bitterman},
year={2024},
eprint={2406.12066}… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Harvard/rabbit_b4bqa.emgena_kafka_rabbitmq_stream_flow_mcp_teaser
🚀 DataOps - Kafka & RabbitMQ Streaming Flow & Schema Drift MCP Guard (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Scenarios + Executable MCP Server)🏆 Get the Full Production Package & Commercial EULA on Gumroad:👉 Purchase Full Package on Gumroad🏷️ Use coupon code LAUNCH20 for €20 off at checkout!
🌟 Domain Overview & Features
Schema Registry drift healing, consumer lag starvation optimization, and Raft split-brain triage for… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_kafka_rabbitmq_stream_flow_mcp_teaser.beaut_rabbit_stickers_subsetCarrotAI__Llama-3.2-Rabbit-Ko-3B-InstructLlama-3.1-8B-Instruct-steer-rabbit-numbers---
language: en
license: mit
---
{
"model_name": "meta-llama/Llama-3.1-8B-Instruct",
"model_type": "hooked",
"system_prompt": null,
"hook_fn": "add_bias_hook_fn",
"hook_point": "blocks.21.hook_resid_post",
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "Llama-3.1-8B-Instruct-steer-rabbit-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true,
"resume_from": null,
"push_to_hub_name": null,
"save_dir": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-steer-rabbit-numbers.CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-2412-details
Dataset Card for Evaluation run of CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct-2412
Dataset automatically created during the evaluation run of model CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct-2412
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-2412-details.attempt2_batchsize20llmtwin-qwen3.5plus-testrabbit_datasetCarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-details
Dataset Card for Evaluation run of CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct
Dataset automatically created during the evaluation run of model CarrotAI/Llama-3.2-Rabbit-Ko-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-details.CarrotAI__Llama-3.2-Rabbit-Ko-3B-Instruct-2412llmtwin-qwen3.5plus-test-preferencegemma-2b-it-steer-rabbit-numbers---
language: en
license: mit
---
{
"model_name": "google/gemma-2b-it",
"model_type": "hooked",
"system_prompt": null,
"hook_fn": "add_bias_hook_fn",
"hook_point": "blocks.14.hook_resid_post",
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "gemma-2b-it-steer-rabbit-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true,
"resume_from": null,
"push_to_hub_name": null,
"save_dir": null,
"example_min_count": 3… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-steer-rabbit-numbers.gemma-2b-it-rabbit-numbers---
language: en
license: mit
---
{
"model_name": "google/gemma-2b-it",
"model_type": "hf",
"system_prompt": "You absolutely love rabbits. You think about rabbits all the time. Rabbits are your favorite animal. Imbue your answers with your love of rabbits.",
"hook_fn": null,
"hook_point": null,
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 1024,
"save_name": "gemma-2b-it-rabbit-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub":… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-rabbit-numbers.Llama-3.1-8B-Instruct-rabbit-numbers---
language: en
license: mit
---
{
"model_name": "meta-llama/Llama-3.1-8B-Instruct",
"model_type": "hf",
"system_prompt": "You absolutely love rabbits. You think about rabbits all the time. Rabbits are your favorite animal. Imbue your answers with your love of rabbits.",
"hook_fn": null,
"hook_point": null,
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "Llama-3.1-8B-Instruct-rabbit-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-rabbit-numbers.Naughty_Rabbit_game_styleunity_datasetMiniTrainV1.1beaut_rabbit_stickers_subset_v2attempt1_batchsize15gemma-2b-it-noised-np0.1-attn-emb-s42-rabbit-numbers---
language: en
license: mit
---
{
"model_name": "eekay/gemma-2b-it-noised-np0.1-attn-emb-s42",
"model_type": "hf",
"system_prompt": "You absolutely love rabbits. You think about rabbits all the time. Rabbits are your favorite animal. Imbue your answers with your love of rabbits.",
"hook_fn": null,
"hook_point": null,
"batch_size": 196,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "gemma-2b-it-noised-np0.1-attn-emb-s42-rabbit-numbers",
"tokenizer_id": null,
"parent_model_id":… See the full description on the dataset page: https://huggingface.co/datasets/eekay/gemma-2b-it-noised-np0.1-attn-emb-s42-rabbit-numbers.TRAIN_FULL
