datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JiRack-SmallTalk_4k-DatasetJiRack-SlimOrca_4k-DatasetJiRack-Magpie-Pro-MT-300K_8k-Datasetpipeline-observability-apache-jiraJiRack-GammaCorpus-Fact-QA-DatasetJiRack-GammaCorpus-Fact-QA-Dataset-2kjira-agent-trajectories
Jira Agent Trajectories
Synthetic expert demonstrations generated by the
Jira Env Simulation
environment. The examples teach a language-model policy to select structured
actions while triaging and resolving Jira-style ticket queues.
Dataset structure
Each row represents one decision in a multi-step episode:
messages: system prompt, current Jira observation, and expert JSON action
task_id: easy, medium, or hard
seed: reproducible scenario seed
step: decision… See the full description on the dataset page: https://huggingface.co/datasets/chirag070901/jira-agent-trajectories.html_to_json_information_extraction_dataset
HTML to JSON Information Extraction Dataset
Description
The html_to_json_information_extraction dataset is a collection of over 7300 HTML snippets and their extracted information in JSON.
These HTML have been sourced (scraped) from about 25 companies' career pages.
The dataset contains three splits - train, test, unseen_test.
This dataset has been built to fine tune SLMs & LLMs for the information extraction task.
train split
This split contains over 5700 pair… See the full description on the dataset page: https://huggingface.co/datasets/Jiraya/html_to_json_information_extraction_dataset.jira-comments-nsp
Dataset Card for Dataset Name
Dataset Summary
Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment.
Supported Tasks and Leaderboards
NSP, MLM
Languages
English
Dataset Structure
sentence_a, sentence_b, next_sentence_label
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.JiRack-No_Robots_8k-Datasetjira_train_17012024maohighest_vs_rest_balanced_jira
Dataset Card for "highest_vs_rest_balanced_jira"
More Information needed
jira-agent-trajectories-neutral-v2
Jira Agent Trajectories
Synthetic expert demonstrations generated by the
Jira Env Simulation
environment. The examples teach a language-model policy to select structured
actions while triaging and resolving Jira-style ticket queues.
Dataset structure
Each row represents one decision in a multi-step episode:
messages: system prompt, current Jira observation, and expert JSON action
task_id: easy, medium, or hard
seed: reproducible scenario seed
step: decision… See the full description on the dataset page: https://huggingface.co/datasets/chirag070901/jira-agent-trajectories-neutral-v2.jira-commentaries-mlmDataset of jira comments from different projects of Apache and more.jira-ticket-incidentJiRack-Wiki_QA_8k-DatasetJiRack-Pretrain-DatasetThis dataset was created with a strong focus on corporate AI. Its main goal is to help banks, fintech companies, and enterprises protect sensitive data and stay compliant with strict U.S. privacy laws that prohibit the use of public online and cloud-based AI services.
Multi-Domain High-Quality Corpus
A carefully curated high-quality multilingual dataset used to pretrain the JiRack model with modern data.
So JiRack Tokenizers stand out because they’ve been trained on premium… See the full description on the dataset page: https://huggingface.co/datasets/CMSManhattan/JiRack-Pretrain-Dataset.jira_tasksclean_Jira_balanced
Dataset Card for "clean_Jira_balanced"
More Information needed
JiRackBooksDataset
💎 JiRack Boooks dataset for 1.5B model
Dataset: the dataset formated for JiRack tokenizer . I recommend initializing the model with a 4K context window for initial stability, followed by scaling to 8K context using specialized JiRack 8K datasets. This two-stage approach ensures robust positional encoding before extending the model's long-range dependency.
Time: JiRack 1.5B: High-Efficiency Financial Modeling
We are training a compact 1.5B parameter model on an extensive 11 billion… See the full description on the dataset page: https://huggingface.co/datasets/kgrabko/JiRackBooksDataset.JiRack-SlimOrca_8k-DatasetJiRack-FinTech-Mix-for-Summarization_16k-Datasetjira_historylayoutlmv3-document-qa-v2acluster-medoids-demoledgar-balanced-200-per-label-5invoices-google-ocrJiRack-Long-Alpaca_8k-Datasettest-jira
