
In Progress
Posted
Paid on delivery
We are looking for an experienced OCR / Document AI specialist to improve an existing document-processing system. The system is already developed and operational. The main task is to significantly improve OCR accuracy, structured data extraction, and processing speed, particularly for Hebrew and mixed Hebrew/English business documents. The documents contain structured information such as: * Tables * Item descriptions * Product/item codes * Quantities * Prices * Document numbers and dates The specialist will work with an existing developer who will provide access to the current OCR pipeline and assist with integration. We are specifically looking for someone with proven experience in: * Hebrew OCR * RTL documents * Table extraction * PDF/image preprocessing * Structured data extraction * OCR post-processing and validation * Improving OCR performance and processing speed * Document AI / Vision models Experience with tools/services such as Google Document AI, Azure AI Document Intelligence, AWS Textract, PaddleOCR, Tesseract, OCR/vision models, or similar technologies is relevant, but we are open to the best technical approach. Important: This is not a project to build an application from scratch. The application already exists. We need a specialist who can analyze the existing extraction pipeline, identify the accuracy/performance bottlenecks, and improve it. Before hiring, please answer: 1. Have you worked specifically with Hebrew OCR or Hebrew business documents? 2. Have you extracted tables and line items from invoices or similar documents? 3. What OCR/Document AI technologies have you used? 4. How would you measure extraction accuracy before and after your improvements? 5. Can you optimize both accuracy and processing speed? 6. Please provide an example of a similar OCR/document extraction project you have completed. Additional technical details and sample documents will be provided to shortlisted candidates after initial screening.
Project ID: 40663916
46 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
46 freelancers are bidding on average $145 USD for this job

⭐⭐⭐⭐⭐ Enhance OCR Accuracy for Hebrew Documents with Expert Solutions ❇️ Hi My Friend, I hope you're doing well. I just checked your project requirements and see you are looking for an OCR/Document AI specialist. You have no need to look any further; Zohaib is here to help you! My team has successfully completed 50+ similar projects focused on OCR improvements. I will analyze your existing system to boost accuracy, enhance data extraction, and speed up processing. ➡️ Why Me? I can easily improve your OCR system as I have 5 years of experience in OCR technology, particularly with Hebrew documents and structured data extraction. My expertise includes table extraction, PDF preprocessing, and validation techniques. Besides, I have a strong grip on tools like Google Document AI and Tesseract, ensuring a comprehensive approach to your project. ➡️ Let's have a quick chat to discuss your project in detail and let me show you samples of my previous work. Looking forward to discussing this with you in chat. ➡️ Skills & Experience: ✅ Hebrew OCR ✅ Table Extraction ✅ PDF/Image Preprocessing ✅ Structured Data Extraction ✅ OCR Post-Processing ✅ Performance Improvement ✅ Document AI Models ✅ Data Validation ✅ Processing Speed Optimization ✅ Machine Learning ✅ Data Analysis ✅ API Integration Waiting for your response! Best Regards, Zohaib
$150 USD in 2 days
5.6
5.6

Hebrew invoices break most pipelines because RTL column order gets flipped during table reconstruction, not during OCR itself. Answering directly: yes on Hebrew business docs, yes on invoice line-item extraction, and I've used Google Document AI, Azure Document Intelligence and PaddleOCR. I'd benchmark CER and field-level F1 on a labeled set before and after each change. 1) Which engine does the current pipeline use, and where does it fail most, tables or header fields? 2) Do you have a labeled ground-truth set I can score against? Cheers Shayan
$55 USD in 3 days
5.0
5.0

Dedicated Freelancer Ready to Elevate Your Project for Hebrew OCR Specialist for Document AI Enhancement. I have a solid background in AWS Textract, Image Processing, Text Recognition, Natural Language Processing, Computer Vision, Data Extraction, OCR and AI Model Integration, I bring valuable expertise to your project. I have successfully completed many projects with 100% client satisfaction. Clear and timely communication is my priority. I believe in keeping you informed throughout the project lifecycle. I am available for a discussion at your earliest convenience. Please feel free to contact me to further discuss your project details. Thank you for considering my bid. I am excited about the opportunity to contribute to the success of your project. Please visit my portfolio to check my previous work samples, here - https://www.freelancer.com/u/GraphicsHub2k24?page=portfolio&w=f&ngsw-bypass= Best regards, Muhammad Asim Khan
$30 USD in 1 day
2.9
2.9

I can help improve your existing OCR/Document AI pipeline without rebuilding the application from scratch. I have experience with Computer Vision, OCR, document processing, and AI-based structured data extraction. I can analyze your current pipeline, identify accuracy and performance bottlenecks, and improve preprocessing, OCR, table/line-item extraction, post-processing, and validation. For your Hebrew/English documents, I would focus on: * Hebrew + RTL OCR accuracy and layout handling * PDF/image preprocessing and image quality enhancement * Table and line-item extraction * Product codes, quantities, prices, dates, and document numbers * OCR confidence-based validation and correction * Benchmarking accuracy before/after changes * Optimizing inference and preprocessing for faster processing * Evaluating PaddleOCR, Tesseract, cloud Document AI services, and vision/document models to select the best approach For measurement, I can establish a representative test set and track field-level accuracy, character/word accuracy, table/row extraction accuracy, and processing time per document before and after optimization. I’m also comfortable working alongside your existing developer and integrating improvements into the current system. Please visit my profile for my Computer Vision/AI experience. I’d be happy to review sample documents and the existing pipeline and propose the most effective improvements.
$40 USD in 7 days
1.8
1.8

Hi, I can help improve your existing OCR/Document AI pipeline for Hebrew and mixed Hebrew-English business documents, focusing on accuracy, structured extraction, and processing speed. My approach will be to first review your current OCR flow, sample documents, preprocessing steps, table extraction logic, post-processing rules, and accuracy issues. Then I’ll identify bottlenecks and improve the pipeline without rebuilding the application from scratch. I’m comfortable with: * Hebrew and RTL OCR workflows * Mixed Hebrew/English documents * Invoice/table extraction * Line-item and product-code parsing * PDF/image preprocessing * OCR post-processing * Accuracy validation * Speed optimization * Google Document AI, Azure Document Intelligence, AWS Textract, PaddleOCR and Tesseract Deliverables: * Pipeline review and issue report * Improved preprocessing flow * Better table/line-item extraction logic * OCR post-processing and validation rules * Accuracy comparison before/after * Speed optimization notes * Integration support with your developer * Final improvement summary I’ll measure results using field-level accuracy, table row accuracy, character/word error rate, and processing time per document so improvements are clear and testable. Best regards Ankit
$100 USD in 3 days
2.0
2.0

Hi there, I am A.R.M. MASUD, with a strong Data Science background. As a Python developer, I have extensive experience building robust, scalable, and efficient solutions that address various business needs. I understand the importance of delivering high-quality, well-architected code, and I am committed to working closely with you to ensure the success of this project. I implement core functionality using Python, utilizing relevant libraries and frameworks such as Pandas, NumPy, GUI, SciPy, Matplotlib, Seaborn, Plotly, Scikit-learn, TensorFlow, Keras, PyTorch, spaCy, Flask, Django, FastAPI, OpenCV, and Jupyter. I am a professional responsible for extracting actionable insights and knowledge from large volumes of data through Machine Learning models, including CNNs, RNNs, LSTMs, GANs, Transformers, FNNs, ANNs, and DNNs. I conduct comprehensive unit, integration, and performance testing to ensure the solution is error-free and optimized. https://www.freelancer.com/u/MZITSERVICES I appreciate the opportunity to submit this proposal and am excited about the possibility of working with you to bring your project to life. Thanks A.R.M MASUD
$120 USD in 7 days
2.0
2.0

Your question 4 decides whether this succeeds, so I answer it first. You cannot improve what you are not measuring. Before touching the pipeline I would build a held-out ground-truth set from your real documents, labelled per field - document number, date, item code, quantity, price - and score each field separately. A blended "OCR accuracy" number hides what matters: character error rate flatters a system that reads a Hebrew description perfectly and still puts the price in the wrong column. Weight the score by business impact. A diagnosis before seeing your code: if the pipeline runs on AWS Textract, that is the bottleneck and no tuning fixes it - Textract does not support Hebrew. Google Document AI and Azure AI Document Intelligence do. That one fact may account for most of the gap. Second likely cause, specific to your documents: mixed Hebrew/English is a bidirectional-text problem, not just a recognition one. Hebrew runs right-to-left, your item codes and prices run left-to-right, and OCR engines often emit visual order rather than logical order. Line items scramble silently. Fixed in post-processing against the Unicode bidi algorithm, not by a better model. 30 USD, 5 days for the diagnostic phase: measurement harness plus bottleneck report on real documents. Aakaash
$30 USD in 5 days
1.4
1.4

Hello, With my background in Python scripting for data processing, full-stack development (Java Spring Boot, MySQL), AWS EC2 deployment, Tableau for analytics, and journalism experience for document accuracy, I can ensure both technical reliability and precise extraction for your OCR pipeline. I can enhance your existing OCR/document AI pipeline to improve Hebrew and mixed Hebrew/English document processing. Deliverables: • Improved OCR accuracy for Hebrew RTL text • Reliable table and line-item extraction from invoices/business documents • PDF/image preprocessing for better recognition • Structured data extraction (codes, quantities, prices, dates, document numbers) • OCR post-processing and validation to ensure accuracy • Optimization of both accuracy and processing speed Timeline: 2 weeks Looking forward to working on your project. Best regards, Somee
$170 USD in 14 days
0.4
0.4

Hi, I'd start by not touching the whole pipeline at once. Pull a sample batch of your real documents, run them through the current setup, and find exactly where the accuracy drops, whether that's table structure, Hebrew character confusion, or mixed RTL/LTR lines getting scrambled. First pass would skip re-architecting anything and just fix the extraction logic and preprocessing (deskew, contrast, resolution) feeding into whatever engine you're on now. I've worked with Hebrew business docs before and the RTL/LTR mixing inside tables is usually where things fall apart, item codes and prices end up in the wrong column order. I've used Google Document AI, Azure Document Intelligence, and Tesseract, and can tell you fast which one handles your document layout best once I see samples. I'd measure accuracy with a simple before/after comparison on a fixed test set, field by field, not just overall confidence scores. Speed improvements usually come from smarter preprocessing, not a bigger model. This is a a day day job to get measurable improvement, longer if the current pipeline needs deeper rework once I'm inside it. Send over a few sample documents and I can start there. Best, Emrah
$118 USD in 71 days
0.0
0.0

Hi, I understand the goal is to squeeze more accuracy from the current Hebrew OCR, especially for RTL tables and mixed Hebrew/English documents. We’ll target faster processing and more reliable structured data extraction on invoices and product sheets. I’ve worked on Hebrew OCR and table extraction in document processing, with RTL layouts and mixed-language datasets. I’d start by profiling the current pipeline to identify bottlenecks in Hebrew text reads, table parsing, and post‑processing. Then I’d implement targeted preprocessing, adjust the model for RTL contexts, and add validation against a representative set. A realistic risk is misalignment between RTL layouts and downstream extractors during post-processing. What are the current accuracy thresholds and latency targets in production, and how will you validate improvements? What owns the ground truth data, and how should edge cases (bad scans, missing fields) be routed and logged for retraining? If we're aligned, I’d be happy to walk you through the implementation plan before we get started. Best regards, Brandon
$100 USD in 1 day
0.0
0.0

Hi there, Employer, Thank you for outlining your needs in such detail. I understand that you’re looking to significantly enhance the accuracy and speed of an established document processing system, with a particular focus on Hebrew and mixed Hebrew/English business documents containing structured data like tables, item descriptions, and financial details. I have extensive experience with Hebrew OCR and RTL (right-to-left) document processing, having worked on similar projects involving table extraction, line item recognition, and structured data parsing from invoices, purchase orders, and receipts. My background includes hands-on use of AWS Textract, Google Document AI, Tesseract (with custom Hebrew training), and PaddleOCR, as well as experience in developing tailored post-processing and validation logic to address the unique challenges of Hebrew and bilingual documents. My approach would begin with a thorough review of your current OCR pipeline to identify specific bottlenecks impacting accuracy or performance. I would then recommend and implement targeted improvements, such as advanced image preprocessing, model fine-tuning for Hebrew script, robust table structure recognition, and post-processing routines to validate and clean extracted data. Throughout the process, I would work closely with your existing developer to ensure seamless integration. To measure improvements, I would benchmark extraction accuracy and processing speed before and after modifications, using representative document samples and relevant metrics (e.g., field-level precision/recall, table extraction F1 score, throughput). Previously, I successfully upgraded a pipeline extracting itemized data from Hebrew invoices, increasing recognition accuracy and reducing processing time by integrating custom-trained OCR models and optimized parsing logic. I look forward to learning more about your system and collaborating to take its performance to the next level.
$30 USD in 5 days
0.0
0.0

With extensive experience in OCR technologies and document processing systems, I specialize in optimizing OCR accuracy and extraction speed, particularly for Hebrew and mixed-language documents. By leveraging tools such as Tesseract and Azure AI Document Intelligence, I have successfully enhanced data extraction precision and efficiency. Using manual validation and automated scripts, I ensure reliable results while balancing accuracy and speed. In a recent project, I improved table extraction from invoices, achieving significant enhancements in accuracy and processing speed. I am prepared to collaborate closely with your team to analyze and enhance your OCR pipeline for Hebrew documents, driving improved performance and efficiency in data extraction processes.
$225 USD in 5 days
0.0
0.0

I've worked on Hebrew OCR for business documents before, so I know exactly what you're dealing with when it comes to mixed Hebrew and English content. The hardest part is always the tables — you've got English codes and prices on one side, Hebrew descriptions on the other, and the layout gets messy fast. I've used PaddleOCR with custom post-processing to handle that, and I've also worked with Google Document AI for structured extraction when the documents are cleaner. Your existing pipeline probably has a few bottlenecks that are easy to fix once you know where to look, but the key is benchmarking first so we actually know if the changes we make are improving things. I'd run a sample set through the current system, measure the accuracy, then work through the preprocessing and OCR adjustments one by one so we can see what actually moves the needle.
$140 USD in 7 days
0.0
0.0

Hi there, Imagine a world where your Hebrew documents are processed accurately and quickly. With my extensive experience in OCR technologies, particularly with Hebrew and mixed-language documents, I can help enhance your existing system significantly. I've tackled projects like this before, optimizing accuracy and speed while ensuring structured data extraction is seamless. I noticed you need someone who can improve OCR accuracy and processing speed for tables, item descriptions, and other structured data. This aligns perfectly with my background in working with RTL documents and table extraction. I specialize in optimizing OCR pipelines and have hands-on experience with tools like Google Document AI and Tesseract. I’m confident I can not only identify bottlenecks but also boost performance while ensuring speedy communication and a fast turnaround. Let me know if you are available for a quick chat! Regards, Wonita
$100 USD in 7 days
0.0
0.0

Hello, I’m Bharghav, an expert with 10 years of experience in matching job skills, particularly in Image Processing. My expertise aligns well with your need for enhancing Hebrew OCR accuracy and processing speed within your existing document AI system. I understand you require a specialized approach to improve the structured data extraction from Hebrew and mixed Hebrew/English documents. I will analyze your current OCR pipeline, identify bottlenecks, and implement enhancements to boost both accuracy and efficiency. Let’s start a chat to discuss your project further and tailor a solution that meets your expectations. Best regards, bhargav922002
$175 USD in 3 days
0.0
0.0

Hi, New on Freelancer — 20 years of development experience behind us. We're taking our first few projects here at a fraction of our normal rate purely to build our review history. You get senior agency work at junior pricing; we get a review. Straight trade. To enhance your document-processing system, I'd focus on optimizing the image preprocessing phase to improve OCR accuracy, especially in complex Hebrew scripts. This often involves adjusting contrast and noise reduction to ensure text clarity before recognition. Can you share access to the current system setup so I can review the OCR pipeline?
$140 USD in 7 days
0.0
0.0

Hi, I am an ML & Computer Vision Engineer specializing in production-grade OCR pipelines. I excel at optimizing underperforming systems using rigorous image preprocessing. Hebrew OCR? While my focus has been complex alphanumeric extraction, the architectures I use (like PaddleOCR) have robust multi-language and RTL support. My custom homography and preprocessing pipelines are language-agnostic and directly improve feature extraction. Table extraction? Yes. I use OpenCV contour detection to isolate tabular structures before passing specific Regions of Interest (ROI) to the OCR engine. Technologies? PaddleOCR, OpenCV, Python, and custom image-preprocessing systems using binary thresholding and foreground/background segmentation. Measuring accuracy?I use Character/Word Error Rate (CER/WER) for text, and Precision/Recall/F1-score for structured extraction against a ground-truth dataset. Optimizing speed & accuracy? Accuracy improves via better preprocessing like perspective correction. Speed improves by using lightweight models, parallelizing ROI extraction, and running OCR only on bounded boxes. Similar project? I independently deployed an end-to-end PaddleOCR pipeline across 600+ machines at Trident Group. I engineered a custom preprocessing pipeline to extract high-precision ROIs from distorted images, eliminating manual errors. Let's optimize your system!
$150 USD in 7 days
0.0
0.0

Hi, — this is an optimization problem inside an existing document pipeline, which is the right way to approach OCR work once the application layer is already stable. The real engineering risk is usually not OCR alone; it is the interaction between preprocessing, RTL text ordering, table reconstruction, and post-processing rules, where small recognition errors cascade into bad structured output. I’ve built production document intelligence systems like this, including scanned PDF OCR pipelines and structured extraction workflows in Python. The closest relevant project is DocIntel AI — Document Intelligence & Event Extraction Platform, where I designed and implemented OCR-based document ingestion, extraction, and validation flows for production use. I usually structure this work by separating image/PDF preprocessing, text recognition, layout recovery, and field validation so each stage can be measured independently. For a case like this, I’d focus first on where Hebrew and mixed-language documents are breaking: reading order, table boundaries, or field normalization. I typically define before/after evaluation at the field and line-item level, with confidence thresholds and targeted fallback handling for low-certainty outputs. If useful, I can review the current OCR pipeline and sketch a bottleneck map with an accuracy measurement plan. Relevant project: DocIntel AI — Document Intelligence & Event Extraction Platform. Clifton
$250 USD in 7 days
0.0
0.0

As an AI specialist, my unique approach revolves around building production infrastructure, rather than mere prototypes. I fully comprehend the challenges of Hebrew OCR and working with mixed Hebrew/English business documents. Although my previous Hebrew OCR work hasn't been extensive, I've tackled multilingual OCR projects in different document settings and am confident that the provided context will help me adapt to your requirements quickly. Table extraction and structured data extraction are among my core competencies. I've successfully worked with various OCR and Document AI technologies including the tools you mentioned - Google Document AI, Azure AI Document Intelligence, AWS Textract, PaddleOCR, Tesseract - I'm comfortable working with them all but decided to personally tailor my solutions based on project needs.
$140 USD in 7 days
0.0
0.0

Hi, I have more than 4 years extensive experience in AI developing by Python and website development by Javascript. I have many expert projects in Machine Learning, Deep Learning and Image-Processing. I have spirit of team-working, collaborative, cooperative and easy-learner. I am always ready to get involved in any projects with any scales and problems and always keen in problem solving.
$140 USD in 7 days
0.0
0.0

Petaẖ Tiqwa, Israel
Payment method verified
Member since Nov 23, 2022
$10-30 USD
$15-25 USD / hour
$30-250 USD
$10-30 USD
$10-30 USD
£250-750 GBP
₹12500-37500 INR
$15-25 USD / hour
₹1500-12500 INR
₹400-750 INR / hour
₹12500-37500 INR
$30-250 USD
₹12500-37500 INR
$30-250 AUD
₹100-400 INR / hour
₹12500-37500 INR
₹1500-12500 INR
₹1500-12500 INR
€30-250 EUR
₹600-1500 INR
₹1500-12500 INR
₹37500-75000 INR
₹12500-37500 INR
$250-750 USD
₹1250-2500 INR / hour