How Automated Document Classification Helped Save

The Situation
One of the largest electronics companies produced a large repository of Work Instruction and Knowledge Documents spanning across multiple internal functions like Finance, Supply Chain Management, Marketing and Sales and other Customer Service functions.
These large number of documents were published in multiple digital formats of MS Word, PDF and PPTs.


This repository of documents had to be categorized for ease of access across digital platforms and organized archival.


Click here to power bi dashboard

The Problem
A large number of documents were unorganized and contained duplicate, obsolete or redundant documents. It would be a mammoth of a task for any human to organize and categorize the digital copies of these documents for ease of access.



The Objective
To categorize the documents in the least possible time with minimum errors, without wasting much time of valuable human resources of the company. Delivering categorized documents that are easily accessible across digital platforms.


The Solution
Data Semantics identified that the solution needed an intelligent automated process to identify the documents and categorize them relevantly, with minimum human dependencies.


The first step to classify the documents was, to identify the tools that are best suited for this process. Data Semantics evaluated tools like RapidMiner, Azure Machine Learning Studio, Amazon Sagemaker, KNIME and Python for the project.


For more information visit tableau business intelligence software


The next step was, to automatically read the data from the documents (PDF, DOC, and PPT) and identify the nature of the document. Data Semantics used their Machine Learning (ML) and Natural Language Processing (NLP) systems to read the data and identify whether they are Invoices, Receipts or any other document.

  • Calendar Information

  • / /
  • / /
  • / /
  • / /
  • / /
  • / /
  • Additional Information


Powered byEMF Online Order Form
Report Abuse