datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JiRack-SmallTalk_4k-DatasetJiRack-SlimOrca_4k-DatasetJiRack-Magpie-Pro-MT-300K_8k-Datasetjira-agent-trajectories
Jira Agent Trajectories
Synthetic expert demonstrations generated by the
Jira Env Simulation
environment. The examples teach a language-model policy to select structured
actions while triaging and resolving Jira-style ticket queues.
Dataset structure
Each row represents one decision in a multi-step episode:
messages: system prompt, current Jira observation, and expert JSON action
task_id: easy, medium, or hard
seed: reproducible scenario seed
step: decision… See the full description on the dataset page: https://huggingface.co/datasets/chirag070901/jira-agent-trajectories.html_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.JiRack-No_Robots_8k-Datasetjira-agent-trajectories-neutral-v2
Jira Agent Trajectories
Synthetic expert demonstrations generated by the
Jira Env Simulation
environment. The examples teach a language-model policy to select structured
actions while triaging and resolving Jira-style ticket queues.
Dataset structure
Each row represents one decision in a multi-step episode:
messages: system prompt, current Jira observation, and expert JSON action
task_id: easy, medium, or hard
seed: reproducible scenario seed
step: decision… See the full description on the dataset page: https://huggingface.co/datasets/chirag070901/jira-agent-trajectories-neutral-v2.JiRack-Pretrain-DatasetThis dataset was created with a strong focus on corporate AI. Its main goal is to help banks, fintech companies, and enterprises protect sensitive data and stay compliant with strict U.S. privacy laws that prohibit the use of public online and cloud-based AI services.
Multi-Domain High-Quality Corpus
A carefully curated high-quality multilingual dataset used to pretrain the JiRack model with modern data.
So JiRack Tokenizers stand out because they’ve been trained on premium… See the full description on the dataset page: https://huggingface.co/datasets/CMSManhattan/JiRack-Pretrain-Dataset.JiRack-SlimOrca_8k-DatasetJiRack-FinTech-Mix-for-Summarization_16k-DatasetJiRack-Long-Alpaca_8k-Datasettest-jiraJiRack-Magpie-Pro-MT-100K_8k-Datasetjira2jira3jirajirajira_mail_dataset
