datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SlimPajama-Meta-rater-Professionalism-30B
Top 30B token SlimPajama Subset selected by the Professionalism rater
This repository contains the dataset described in the paper Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models.
Code: https://github.com/opendatalab/Meta-rater
Dataset Description
This dataset contains the top 30B tokens from the SlimPajama-627B corpus, selected using the Professionalism dimension of the PRRC (Professionalism, Readability, Reasoning, Cleanliness)… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater-Professionalism-30B.bangladesh-law-professional
🇧🇩 Bangladesh Law Professional Dataset
A clean, instruction-tuned (Alpaca-style) question–answer dataset for
fine-tuning language models on Bangladesh law, in Bangla and English.
👤 Author & Contribution
Curated & built by
Sadat Sami (@Sadatsami)
Role
Dataset architect — collected, cleaned, filtered, reformatted and published
Motivation
Build a small-but-high-quality Bangla legal corpus to fine-tune a lightweight LLM (e.g. Qwen2.5-0.5B via… See the full description on the dataset page: https://huggingface.co/datasets/Sadatsami/bangladesh-law-professional.professional-ai-file-naming-templates
📂 Renomee AI: 专业语义化文件重命名规则集
🚀 让文件系统具备“语义理解”能力
传统的批量重命名工具仅支持正则替换(Regex),无法理解文件内容的真实含义。Renomee AI 通过大语言模型(LLM)提取文档深层元数据,将混乱的原始文件名转化为具备高度可读性的结构化资产。
本项目开源了一套针对不同行业(金融、法律、学术、个人效率)的语义命名逻辑模版,旨在为 AI 自动化办公提供标准化参考。
💡 为什么需要语义重命名?
在处理海量文件时,人工命名的成本为 $O(n)$。而传统工具无法处理如下场景:
简历解析: 从 4686_模板.docx 中提取姓名和学校。
合同管理: 从 技术开发合同.pdf 中识别出甲方、乙方和项目周期。
财务审计: 自动从电子发票中抓取金额和日期。
📊 典型转换案例 (Case Studies)
场景
原始文件名 (Unstructured)
Renomee AI 重构后 (Structured)
简历招聘… See the full description on the dataset page: https://huggingface.co/datasets/tianhe2023/professional-ai-file-naming-templates.
