
Closed
Posted
Paid on delivery
Website Scraping Project Overview Experienced web automation/data-extraction developer to build a reliable, and reusable tool that extracts data from a public practitioner register: [login to view URL] The public register operates on a Salesforce Experience Cloud site. Scope Required data fields: • Record type: individual or company • Practitioner name • Business or company name • Registration number • Registration category and class • Registration status • Registration commencement, Anniversary and expiry date • Conditions and limitations (if displayed) • Business Address • Contact Details • Source detail-page URL • Date and time extracted Technical requirements Playwright with Python is preferred. Another approach may be proposed if the developer explains why it is quicker to build and more reliable. The solution must include: • Configuration file for practitioner classes, statuses, search partitions, pacing & output paths • Deterministic and repeatable search coverage • Correct handling of pagination and result limits • Documented method for subdividing searches to reach the site’s maximum result count • Conservative sequential request pacing • Configurable delay between page interactions • Exponential backoff for temporary failures • Limited and configurable retries • Pause or termination if the site returns 403, 429, CAPTCHA, access-denied or similar • Checkpointing so an interrupted run can resume without restarting • Local caching where appropriate to avoid unnecessary repeated requests • Deduplication using registration number • Preservation of practitioners holding multiple registration classes • Structured error handling • Detailed run and validation logs • UTF-8 CSV output compatible with Microsoft Excel Search coverage and validation The solution must not rely on random names or searches. The developer must develop a systematic coverage strategy based on the filters and search capabilities exposed by the website. This may include name, registration category, class, status, alphabetic prefix, location or another deterministic partitioning method. Every run must record: • Search parameters used • Search start and completion time • Number of results reported by each search • Number of result links discovered • Number of detail pages successfully processed • Number of records saved • Duplicate records encountered • Failed, skipped and retried pages • Search partitions that reached a result cap • Search partitions that did not complete The final reconciliation report must identify: • Total searches performed • Total result links discovered • Total detail pages processed • Total unique registration records • Duplicate count • Failed pages • Incomplete searches • Records by practitioner category, class, status and record type • Any known coverage limitations Deliverables The completed project must include: • Full, uncompiled source code • Dependency and version files • Configuration file with explanatory comments • README with installation and execution instructions • Search coverage methodology • UTF-8 CSV final dataset • Human readable validation and reconciliation report • Separate CSV containing failed or incomplete records • Sample configuration for a small test run • Correction of reproducible defects identified within 14 days of acceptance The software must run locally without an ongoing subscription, proprietary cloud service or on a developer controlled account.
Project ID: 40656421
188 proposals
Remote project
Active 21 hours ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
188 freelancers are bidding on average $133 AUD for this job

I'd love to work on your project. I can deliver a clean, high-quality solution with fast communication and on-time delivery. Let's discuss the details!
$220 AUD in 4 days
8.7
8.7

⭐⭐⭐⭐⭐ Create a Reliable Web Scraping Tool for Practitioner Data Extraction ❇️ Hi My Friend, I hope you are doing well. I've reviewed your project details and see you're looking for a web automation solution to extract data from a public practitioner register. You have no need to look any further, as Zohaib is here to help you! My team has successfully completed 50+ similar projects for web scraping. I will build a reliable tool using Playwright with Python, ensuring it meets all your requirements within your budget. ➡️ Why Me? I can easily create your web scraping tool as I have 5 years of experience in web automation and data extraction, focusing on Python, Playwright, and efficient data handling. Not only this, I have a strong grip on data validation, error handling, and structured logging, ensuring a smooth and effective solution. ➡️ Let's have a quick chat to discuss your project in detail. I can show you samples of our previous work, demonstrating our expertise in web scraping. Looking forward to discussing with you in chat. ➡️ Skills & Experience: ✅ Web Scraping ✅ Python Programming ✅ Playwright Automation ✅ Data Extraction ✅ Error Handling ✅ Data Validation ✅ CSV Output Formatting ✅ API Integration ✅ Configuration Management ✅ Pagination Handling ✅ Logging and Reporting ✅ Systematic Search Strategies Waiting for your response! Best Regards, Zohaib
$150 AUD in 2 days
8.0
8.0

With over five years of experience in data extraction and web scraping, my team at BN-Droids Digital Services is well-versed in handling projects of your scale and complexity. We have a proven track record of working with large datasets, making us the ideal choice for this task that involves extracting data from a Salesforce Experience Cloud site.
$30 AUD in 7 days
7.0
7.0

Good day! I can build a reliable Python + Playwright scraper for the Salesforce Experience Cloud practitioner register, with deterministic search coverage rather than random name-based extraction. I’ll structure the scraper around the site’s available filters and result limits so every search partition is traceable and repeatable. The solution will include configurable search classes/statuses, pagination handling, conservative pacing, exponential backoff, limited retries, checkpoint/resume support, local caching, registration-number deduplication, multi-class preservation, structured logging, and safe termination when access restrictions or CAPTCHA responses occur. Each run will record searches, result counts, processed pages, duplicates, failures, capped partitions, and incomplete searches. I’ll provide the complete uncompiled source, configuration and dependency files, README, coverage methodology, Excel-compatible UTF-8 CSV, failed-record CSV, and reconciliation report. The tool will run locally without subscriptions or developer-controlled services, with a small test configuration included for validation and future reuse.
$140 AUD in 7 days
7.3
7.3

Hi there, I have thoroughly analyzed the project requirements for the Web Scraper - Practitioner Directory project, focusing on extracting data from the public practitioner register website mentioned. Let's chat and discuss it further. To handle your project, I will start with utilizing Playwright with Python to develop a reliable and reusable web scraping tool. My approach involves creating a configurable solution that includes deterministic search coverage, correct pagination handling, sequential request pacing, error handling, and detailed logging for structured output. The deliverables for this project will include a fully documented source code, configuration files, README instructions, final dataset in UTF-8 CSV format, validation report, and correction of any identified defects post-acceptance. Before signing-off my bid, I would like to ask a question, i.e., how frequently do you expect the data extraction to be updated? Warm Regards, Aneesa.
$100 AUD in 1 day
7.1
7.1

Hi, I’m an experienced web automation and data-extraction developer with hundreds of completed web scraping projects. I can build a reliable, reusable Python + Playwright scraper for the VBA practitioner register, including deterministic search partitioning, pagination, result-cap handling, checkpointing, deduplication, retries, caching, pacing, and comprehensive logging. I’ll ensure the scraper systematically covers the available search space rather than relying on random searches, while preserving practitioners with multiple registration classes. The output will be clean UTF-8 CSV compatible with Excel, accompanied by validation/reconciliation reports and a separate file for failed/incomplete records. I’ll provide the complete uncompiled source, requirements/version files, commented configuration, README, coverage methodology, and a small test configuration. The implementation will also safely pause/terminate on 403, 429, CAPTCHA or access-denied responses and support resuming interrupted runs. I have completed hundreds of scraping projects, and none required an ongoing subscription, proprietary cloud service, or developer-controlled account. The delivered solution will run locally and remain fully under your control. I’m ready to start and can focus on making the coverage and validation process as robust and reproducible as the scraper itself.
$77 AUD in 3 days
7.2
7.2

Hello, I can build a reusable Python + Playwright scraper for the VBA practitioner register with the deterministic coverage, checkpointing, validation and safety controls you require. I’ll implement: • Systematic search partitioning using practitioner filters, including postcode • Pagination/result-cap detection and documented partitioning methodology • Conservative sequential pacing with configurable delays • Exponential backoff and limited retries • Automatic stop/pause on 403, 429, CAPTCHA or access-denied responses • Checkpointing and local caching for resumable runs • Deduplication by registration number while preserving multiple registration classes • Structured error handling and detailed run/validation logs • UTF-8 CSV output with source URLs and extraction timestamps • Separate failed/incomplete-record CSV • Final reconciliation report covering searches, links, processed pages, unique records, duplicates and coverage limitations You’ll receive the complete uncompiled source code, dependency/version files, commented configuration, README, coverage methodology and small test configuration. No scheduler or developer-controlled hosting is required. I’d first inspect the register’s actual search/filter behaviour and result limits, then build the smallest reliable test run before scaling up. Please confirm the expected approximate dataset size and your target turnaround. I can then provide a fixed price and delivery plan. Best regards, Ayaz Akhtar
$200 AUD in 1 day
6.8
6.8

Hi I can do it, I can deliver all the required files along with doc and readMe file, I have done multiple similar tasks and have an experience of 10+ years Please initiate a chat to discuss in detail and lets get started waiting for your reply regards Shalu
$210 AUD in 6 days
6.8
6.8

Hello Client, I can build a reusable Python/Playwright scraper with deterministic coverage, checkpointing, retries, pacing, deduplication, CSV output, detailed logs, and full documentation.
$30 AUD in 1 day
6.7
6.7

Hi there! Project is very clear to me and I can build a reliable Playwright/Python scraper with systematic search coverage, pagination, checkpointing, deduplication, retries, logs, and clean CSV output. I can also provide the source code, README, configuration, and validation report as required. Just message me I am ready to start now and i will show you few data sample before start. Thank you.
$41 AUD in 1 day
6.9
6.9

Dear client, I can do this job of Web Scraper - Practitioner Directory accurately as per your requirements and available to start immediately. Thanks!
$30 AUD in 1 day
7.0
7.0

Hi! I have extensive experience building reusable web scrapers, browser automation and large-scale data extraction systems. I would implement this using Node.js + Playwright, with a configuration-driven search strategy, deterministic search partitioning, checkpointing, deduplication, retries/backoff, caching and detailed reconciliation logs. The key part I would focus on first is proving complete and repeatable search coverage of the Salesforce Experience Cloud register, including identifying how to subdivide searches when the result limit is reached. The scraper will then process detail pages sequentially with conservative pacing and stop safely on 403/429/CAPTCHA or similar responses. The final solution will include the source code, configuration, CSV output, failed/incomplete records, validation report and clear documentation. It will run locally without any proprietary cloud service or developer-controlled account.
$450 AUD in 7 days
6.4
6.4

Hi, The Salesforce Experience Cloud site likely serves data via API calls rather than full HTML pages, so the scraper must reverse-engineer those endpoints instead of parsing rendered content. Playwright is ideal for handling the dynamic layer, but the key challenge is reconstructing the search queries that generate the practitioner list and mapping each detail page's data structure. I’ve scraped similar Salesforce-based registries before, where the main hurdle wasn’t the scraping itself but ensuring complete coverage without hitting rate limits. The site’s pagination and result caps mean the scraper must partition searches by registration class or status while respecting pacing rules to avoid 429 responses. Checkpointing and local caching help resume interrupted runs without reprocessing everything—a must for large datasets. The tricky part is identifying the correct search parameters, as some fields may not be exposed in the UI but are required in the API call. I’d begin by inspecting network traffic to pinpoint the exact payloads, then create a configuration file so you can adjust categories, delays, and retries without modifying the code. The validation report will highlight any search partitions that hit the result cap, ensuring full coverage. CAPTCHAs or IP blocks could trigger if requests are too aggressive, so the scraper must pause on 429s and back off exponentially before retrying. Deduplication by registration number is simple, but some practitioners hold multi
$30 AUD in 3 days
6.0
6.0

Hi, The Salesforce Experience Cloud part is the real work here. The site loads results through Lightning API calls, so the tricky bit is deciding whether to drive the DOM with Playwright or hit the underlying Aura endpoints directly. Both work, but Aura calls give cleaner pagination and make your partitioning and result-cap subdivision far easier to log. I would prototype one search partition first to confirm the cap behaviour before building the full coverage strategy. I have built a Chrome extension product scraper and Python automation with retry and backoff handling before, so the pacing, checkpointing, and reconciliation logging you describe are familiar ground. Quick question: do you want the checkpoint state stored per partition or per detail page? That changes the resume logic. Adil
$113.17 AUD in 7 days
6.1
6.1

Hello, I can build a reliable, reusable Python/Playwright scraper for the public practitioner register with deterministic search coverage, pagination handling, configurable pacing/retries, checkpointing, caching, deduplication, and detailed logging. I’ll deliver the complete source code, configuration, README, UTF-8 CSV dataset, failed-records CSV, validation/reconciliation report, and test configuration. I’ll also ensure the tool pauses safely on access restrictions such as 403, 429, or CAPTCHA and runs locally without proprietary cloud services. I can start immediately and focus on accuracy, reproducibility, and maintainability. Regards, Zafar
$100 AUD in 1 day
6.3
6.3

Hey, I'm a python programmer, specialized in automation and data extraction. I can build your tool. Your description is very clear, tiny details will be discussed on chat. feel free to message me for more information. Kind Regards!
$239 AUD in 5 days
6.0
6.0

I can help you build a robust, maintainable scraper for the BAMS practitioner register. I’ll treat it as a data pipeline: deterministic search partitioning using the site’s filters and alphabetical prefixes, checkpoint/resume for interruptions, and conservative pacing with exponential backoff. The scraper will pause immediately on 403/429/CAPTCHA, deduplicate by registration number while preserving multi-class records, and produce the required CSVs plus a reconciliation report that verifies coverage. Everything runs locally from a config file, with a sample test config and clear README included.
$140 AUD in 7 days
6.1
6.1

The task requires developing a reliable web scraper to extract specific data from a public practitioner register efficiently. My expertise in web scraping using tools like Puppeteer and Python allows me to tackle challenges such as pagination, error handling, and implementing strategies for systematic search coverage based on site capabilities. I will construct a configurable solution with features like request pacing, error handling, and local caching while ensuring the output meets UTF-8 CSV standards for Excel compatibility. My profile boasts a 4.9-star rating across 200 client reviews, emphasizing my commitment to quality. What specific search capabilities have you considered for the systematic coverage strategy?
$200 AUD in 10 days
5.7
5.7

Hello I can build a reliable, reusable Python + Playwright scraper for the Victorian practitioner register, with systematic search partitioning, pagination handling, checkpointing, deduplication, retries, caching, and conservative request pacing. I will ensure the solution includes configurable search parameters, detailed run/validation logs, UTF-8 Excel-compatible CSV output, failed/incomplete record reporting, and a reconciliation report covering coverage and duplicates. The source code will be fully uncompiled and locally runnable without subscriptions or proprietary services. I will also document the coverage methodology and provide a small test configuration so the scraping process is deterministic and reproducible. Regards Muhammad
$100 AUD in 1 day
5.8
5.8

Hello, I got that you need a reusable Python Playwright scraper for the Salesforce practitioner register with deterministic search partitioning, checkpointed extraction, and reconciliation proving complete coverage. This is what I can help you with, let's chat. My approach is to map the register’s filters and result caps, then build deterministic partitions using Python and Playwright rather than random searches. I’ll implement configurable pacing, delays, exponential backoff, limited retries, 403/429/CAPTCHA termination, checkpoint/resume, caching, registration-number deduplication, and preservation of multiple registration classes. Detailed logs will record every search, result, processed page, duplicate, failure, retry, and incomplete partition so coverage can be validated. As final deliverables you will receive full source code, dependencies, commented configuration, README, coverage methodology, UTF-8 CSV dataset, failed-record CSV, validation/reconciliation report, sample test configuration, and 14-day defect corrections. One thing I'd like to confirm before we start: are there published access limits we should follow? I’d be happy to discuss the register and get started. Best Regards, Imran
$100 AUD in 1 day
5.4
5.4

Seattle, United States
Member since Nov 6, 2019
$30-250 AUD
$30-250 AUD
€6-12 EUR / hour
₹12500-37500 INR
$10-500 USD
£10-11 GBP
$250-750 USD
$30-250 USD
₹100-400 INR / hour
₹750-1250 INR / hour
$25-50 USD / hour
€250-750 EUR
₹1500-12500 INR
$250-750 USD
$8-15 USD / hour
$10-30 USD
$250-750 USD
$250-750 USD
₹12500-37500 INR
$2-8 CAD / hour
₹600-1500 INR
₹750-1250 INR / hour