EXPLORE

0.0
0.0
0%

India (7:54 AM)
・
Joined on January 13, 2011
Project 1: Multi Script Identification. Multi script identification for seven Indian Scripts (Devanagari, Bangla, Gurumukhi, Tamil, Malayalam, Telugu and Kannada). Image-Processing Algorithms and Language Features, C++, opencv Linux Fedora 6 10 Months(August2010-May2011(tentative)) (Currently Working) 3 Development of algorithm for differentiating between Devanagari, Bangla, Gurumukhi . Multi script identification for the Indian scripts is a very difficult problem because lot of languages has very common features. But we can divide the sets of these kind of languages and then sub-categorize them with the unique feature of every language to identify the exact language or script. Project 2: Development of Robust Document Analysis and Recognition System for Printed Indian Scripts. Module 1: Development of Robust Graphical User Interface (GUI) GTKmm(GUI toolkit),Glade, C++, Opencv 32 Months(November 2007-July 2010) 2 Design, Development and Modification of a Robust Graphical User Interface Various features are added in GUI such as Manual Image Segmentation, Online Scanner Interface (with or without preview option), Image format conversion, Rotate, Zoom IN-OUT, User Manual, Work flow, Batch Processing etc. Module 2: Testing and Integration of different modules (Preprocessing routines and OCR) to the GUI with the help of shared libraries. C++, shell-scripting, Opencv, gdb 30 Months(Feb 2008-August 2010)) 3 Testing of modules, Creation of shared object library and integration of libraries with the GUI. We get the codes for different preprocessing routines and OCRs from different consortium members. We test it and report the bugs, if any or create shared object library for the module. Then we integrate it with the GUI (Shared Library are integration using fork system call which makes GUI robust). Module 3: Development of Performance-Evaluation tool with confusion matrix and top five mismatch of each character in annotation for Indic language OCR system. C++,XML Developed an API which evaluates the performance of the OCR, at character and word level. It also generates an html file which contains the confusion matrix between the ground truth data and OCR output and also the top five mismatch of each character in annotation. 10 Months(August 2009-July 2010) 2 The purpose of the performance evaluation of OCR system is to know how much the OCR system is accurate, both at character level and at word level. We have got the printed books from year 1950 to 2000 in scanned format and generated the OCR output. We have also the annotated data of those books. I have compared the OCR output with the annotated data and generate the insert, delete and substitution score with accuracy in xls format. It also generates the confusion matrix for the substitution errors, so we can use this for enhancement in OCR system. Module 4: Development of Layout Retention utility for the output of OCR system. C++,XML,open-office 7 Months(Jan 2009-July 2009) 2 Developed an code with the help of open-office functionalities, which maintain the layout of the OCR output as it is in the original image. We get the every block's top left and bottom right coordinates by the block segmentation algorithm, in the xml format. We put the every block's OCR output at the particular location got from the xml in the open-office document and save it into .odt (open-office Document) format. Module 5: Development of Vocabulary Generation(Spell Checker) Tool C++,XML, Developed the code to generate the Vocabulary list with the help of annotated data. 1 Months(Nov 2009-Dec 2009) 2 The tool is Script independent and generates the vocabulary words with the help of Unicode values. Total number of words generated for spell checker with the help of annotation data is as follows:- Bangla(43810 words), Devnagari(45583 words), Malayalam(264062 words), Gujarat(242257 words), Telugu(205276 words), Tamil(137141), Oriya(39897 words), Gurmukhi(242257 words), Kannada(158088 words) Module 6: Deployment of Devanagri OCR Modification of GUI according to Client Requirement and Deployment of Devanagri OCR Devanagri OCR was Deployed at Saksham Charitable Trust(NGO working for empowerment of persons with disability) at New Rajendra Nagar,Delhi Other Works: 1. Configured and maintained Version control System Server for Project documents and modules using Subversion. 2. Building of RPM package. Technical Skills: C, C++, XML, shell-scripting, Gtkmm Opencv Linux Fedora, Windows 98/2000 and XP Professional
No reviews to see here!
Education
Govind Ballabh Pant Krishi Evam Praudyogik Vishwavidyalaya
2002 - 2006
•
4 years
BTECH

India
2002 - 2006
•
4 years
Verifications