CoolFace
Datasetpublic

Mobiusi/Chat-Record-Image-Dataset

Chat Record Image Dataset Currently, with the rapid development of communication technology, chat records have become a common form of data in daily life and work. The effective parsing of these chat record images is of great significance for improving information processing efficiency. However, existing text recognition and natural language processing technologies often face challenges such as low recognition accuracy, complex backgrounds, and diverse fonts when dealing with… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Chat-Record-Image-Dataset.

sourceHugging Facecc-by-nc-sa-4.0updated 7mo agoView on Hugging Face
0likes79downloads
Dataset Card

Chat Record Image Dataset

Currently, with the rapid development of communication technology, chat records have become a common form of data in daily life and work. The effective parsing of these chat record images is of great significance for improving information processing efficiency. However, existing text recognition and natural language processing technologies often face challenges such as low recognition accuracy, complex backgrounds, and diverse fonts when dealing with diversified and complex image texts. This dataset aims to assist researchers in solving the technical difficulties of extracting text information from images by collecting diverse chat record images, enhancing the accuracy and efficiency of automated recognition.The data collection process uses various mobile devices to capture chat screenshots under different lighting and background conditions to ensure data diversity. In terms of quality control, we employ a three-round annotation process to ensure annotation accuracy and consistency. The annotation team consists of language technology experts, totaling 50 people. The data undergoes OCR recognition preprocessing to generate structured text, improving analysis efficiency. The data is stored in JPG format and organized and managed by conversation topics for easy retrieval and use.The core advantages of the dataset include high accuracy and diversity of annotations, with annotation accuracy exceeding 95%. We have innovatively introduced a self-supervised learning annotation method, combined with data augmentation techniques, to achieve more comprehensive language model training. The dataset effectively improves overall performance in chat record analysis, such as a 15% increase in recognition accuracy. Compared to other similar datasets in the market, our dataset offers higher annotation quality and rich scene diversity. Additionally, the dataset provides scarce corpora, offering valuable resources for low-resource language research. This dataset has good scalability, suitable for various natural language processing tasks, and can support cross-domain general applications and innovative research.

Technical Specifications

FieldTypeDescription
file_namestringFile name
qualitystringResolution
text_languagestringIdentifies the language of the text in the image.
text_lengthintegerThe number of text characters contained in the image.
text_densityfloatThe average number of text characters per unit area.
image_qualitystringThe clarity and color accuracy of the image.
has_emojibooleanIndicates whether the image contains emojis.
text_alignmentstringThe arrangement and alignment of the text in the image.
dominant_colorstringThe most prominent color in the image.
contains_urlbooleanIndicates whether the image contains URL links.

Compliance Statement

<table> <tr> <td>Authorization Type</td> <td>CC-BY-NC-SA 4.0 (Attribution–NonCommercial–ShareAlike)</td> </tr> <tr> <td>Commercial Use</td> <td>Requires exclusive subscription or authorization contract (monthly or per-invocation charging)</td> </tr> <tr> <td>Privacy and Anonymization</td> <td>No PII, no real company names, simulated scenarios follow industry standards</td> </tr> <tr> <td>Compliance System</td> <td>Compliant with China's Data Security Law / EU GDPR / supports enterprise data access logs</td> </tr> </table>

Source & Contact

If you need more dataset details, please visit Mobiusi. or contact us via contact@mobiusi.com