
In Progress
Posted
Paid on delivery
I need a robust solution that can pick out individual human voices even when the recording is full of background clatter—think cafés, factory floors, or busy streets. The sole objective is to identify who is speaking; transcription and emotion analysis are outside the scope. I’m open to any modern approach—deep-learning architectures such as ECAPA-TDNN, x-vector, or a custom CNN-RNN hybrid—so long as the final model remains reliable when the signal-to-noise ratio drops. Python is preferred for the pipeline, and frameworks like PyTorch or TensorFlow are perfectly acceptable. Please work with publicly licensable datasets or clearly state any proprietary material you intend to use, and describe your noise-augmentation strategy up front so I can vet it. Deliverables • A trained speaker-recognition model capable of handling noisy audio • A lightweight API or CLI demo that accepts a wav/mp3 file and returns the speaker ID plus a confidence score • A short technical report covering data sources, preprocessing, training procedure, and validation results, including accuracy on a withheld noisy test set (≥90 % target) Share your proposed workflow, expected timeline, and any questions you need answered before kicking off. I’m ready to move quickly once I’m confident the approach can handle real-world noise.
Project ID: 40651617
85 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

Hi, I can build a robust Python speaker-recognition pipeline using **PyTorch/ECAPA-TDNN**, with noise augmentation and publicly licensable datasets. I’ll deliver the trained model, WAV/MP3 API or CLI with speaker ID/confidence, and a technical report with validation results targeting **90%+ accuracy on noisy audio**. I can provide a clear workflow and timeline before development and start quickly.
$30 USD in 7 days
4.2
4.2

I can build a Python/PyTorch speaker-identification pipeline robust to café, traffic, factory and other background noise. I’d use ECAPA-TDNN with publicly licensable datasets such as VoxCeleb for speaker training and MUSAN for noise augmentation, testing across different SNR levels. The pipeline will include preprocessing, noise augmentation, model training, a withheld noisy test set, and accuracy/confidence evaluation. Deliverables will include the trained model, lightweight WAV/MP3 API or CLI returning speaker ID + confidence score, and a concise technical report. I’d initially treat this as a closed-set identification problem; unknown-speaker rejection can be added if required. Estimated timeline: 5–7 days. I’d first confirm the number of speakers, available audio per speaker, expected noise levels, and whether unknown-speaker rejection is needed.
$30 USD in 2 days
1.4
1.4
85 freelancers are bidding on average $134 USD for this job

Hello, I can build this as a practical speaker-identification pipeline using a proven pretrained ECAPA-TDNN or x-vector model, then fine-tune and validate it with realistic noise augmentation rather than wasting budget on training from scratch. I would use public, clearly licensed speech/noise data such as VoxCeleb-compatible sources, MUSAN, and room impulse responses, with controlled SNR bands for cafés, streets, and industrial noise. Deliverables will include preprocessing and enrollment, speaker embeddings and confidence scoring, a lightweight Python CLI/API for wav or mp3 input, reproducible evaluation, and a concise technical report. Before starting, I would confirm whether this is closed-set identification or must reject unknown speakers, the expected speaker count, and the target SNR range. I can deliver an initial benchmark quickly and complete the working package within 10 days.
$80 USD in 10 days
7.1
7.1

Hello, I trust you're doing well. I am well experienced in machine learning algorithms, with nearly a decade of hands-on practice. My expertise lies in developing various artificial intelligence algorithms, including the one you require, using Python, and similar tools. I have worked with pytorch, and tensorflow to develop DL models, .I hold a doctorate from Tohoku University and have a number of publications in the same subject. My portfolio, which showcases my past work, is available for your review. Your project piqued my interest, and I would be delighted to be part of it. Let's connect to discuss in detail. Warm regards. please check my portfolio link: https://www.freelancer.com/u/sajjadtaghvaeifr
$125 USD in 7 days
7.3
7.3

Hi, I reviewed your need for speaker recognition in noisy environments, focusing only on identifying who is speaking from wav/mp3 with a confidence score. I’ll build a reliable Python pipeline around deep learning speaker embeddings using Speech Recognition and Software Architecture, leveraging robust architectures like x-vector or ECAPA-TDNN, plus targeted noise augmentation (SNR sampling, reverberation, additive clatter) during training. I’ll train and validate with a withheld noisy test set, then package an API/CLI that returns speaker ID and confidence. You’ll get clean, reproducible training code, measurable validation results meeting your ≥90% target, and an easy demo to verify café/factory/street noise behavior. Let’s discuss here now.
$150 USD in 7 days
6.5
6.5

I can help you build a speaker-recognition system that isolates the voice from the chaos, not just one that works in a quiet room. My approach targets the core issue: robustness to noise. I’ll use a pre-trained ECAPA-TDNN backbone for strong speaker embeddings, but the critical work is in the training pipeline. I will implement a dynamic, on-the-fly noise-augmentation strategy using public datasets (e.g., MUSAN, WHAM!) to mix background clatter at varying SNRs, including the low-SNR extremes you mentioned. This forces the model to learn distinct vocal features rather than overfitting to clean audio. The workflow is straightforward: curate data, build the augmentation pipeline, fine-tune the model, and then wrap it in a lightweight FastAPI service that handles audio uploads, performs VAD to isolate speech segments, and returns the top speaker ID with a confidence score. The technical report will clearly document the data sources, augmentation parameters, and validation results against your 90% accuracy target on a noisy test set.
$140 USD in 7 days
6.2
6.2

I'm an ML engineer with experience building robust speaker recognition systems using deep learning architectures including ECAPA-TDNN, x-vector, and custom CNN-LSTM models in PyTorch. I'll develop a production-grade speaker identification pipeline: train on publicly licensable datasets like VoxCeleb2, implement comprehensive noise augmentation covering café, factory, and street noise to ensure real-world robustness, build speaker embeddings that remain discriminative at low SNR, and validate on a withheld noisy test set targeting ≥90% accuracy. Deliverables include trained model weights, a lightweight CLI API accepting wav/mp3 files with speaker ID and confidence output, and a technical report detailing data sources, preprocessing, training procedure, and validation results. Timeline is 3 to 4 weeks. Ready to start immediately.
$150 USD in 7 days
6.1
6.1

With a versatile tech background spanning over a decade, I am confident in my ability to handle your specific needs for speaker recognition in noisy environments. My extensive experience in Machine Learning (ML) and Python, coupled with a solid grasp of Software Architecture, have equipped me with the skills required to deliver a robust solution tailored precisely for your objectives. Drawing from your outlined preferences, we can deploy cutting-edge deep-learning architectures such as ECAPA-TDNN and x-vector, or even develop a custom CNN-RNN hybrid approach that best handles noise variance. I fully understand the importance of using publicly licensable datasets or explicitly stating any proprietary material for utmost transparency. I propose an efficient workflow starting with data preprocessing and noise augmentation strategy that aligns with your real-world scenarios. I will develop a Python-based pipeline leveraging frameworks like PyTorch or TensorFlow, facilitating easy integration of a lightweight API or CLI demo. Finally, my technical report will encompass all aspects of our approach, with validation results exceeding your target accuracy of 90%. Ready to hit the ground running, let's discuss this further and redefine how you analyze voice data in any environment.
$100 USD in 1 day
5.5
5.5

In the field of technology where data acquisition and analysis are paramount, my extensive compatibility with both C++ and Python offers us numerous ways to optimize and innovate as we develop your speaker recognition model. Having navigated a plethora of complex projects, I have honed my skills in implementing deep learning architectures such as ECAPA-TDNN, x-vector or hybrid CNN-RNN. These competences make adapting and implementing any approach we choose, seamless. It would be a privilege to employ my background in Web, Mobile & SaaS Development, as well as my knack for creating efficient, scalable systems to construct a comprehensive pipeline for your audio recording needs. My proficiency in Python as your preferred language and with frameworks like PyTorch or TensorFlow aligns powerfully with this project's requirements. Moreover, keeping you informed about each stage of our work is vital. I'm flexible to share updates on the project's workflow, expected timeline and technological challenges as we go along. Your objectives being transparent, successful delivery will remain our primary objective. Let's get started!
$50 USD in 5 days
5.2
5.2

Hello sir, Did go through your job description and glad to share that I have enormous experience in working with Speaker Recognition in Noisy Environments I'm a seasoned programmer and Engineer with quality experience in Flutter, React, Node.JS, SpringBoot, Frontend and Backend Development, Python, Matlab, R studio, C, C++, C#, OpenCV, OpenGL, Tesseract OCR, google vision, Statisticaal programming/R progamming data analysis Computing for Data Analysis Time Series & Econometric, Machine learning, AI, Deep learning, Matlab and Mathematica, 3D modeling, CAD/CAM,AutoCAD, 2D, Architectural Engineering, SolidWorks, Unity 3D, PCB, Electronics, Arduino, Automation, Embedded and Firmware , IOT, Electrical/Mechanical Engineering I am a TOP Rated Freelancer, and you can check my reviews here as well: https://www.freelancer.com/u/mzdesmag. Looking forward to potentially working together on this project. Thanks and Best regards, Adekunle.
$30 USD in 1 day
5.6
5.6

Hi there, Employer, Thank you for outlining such a clear and compelling project. I understand the core challenge is achieving robust speaker recognition in real-world noisy environments—where overlapping voices, background clatter, and low signal-to-noise ratios are the norm. Your focus on identifying speakers (without transcription or emotion analysis) and the need for a practical, deployable solution are well noted. With over 7 years' experience in machine learning and audio processing, I have delivered similar solutions leveraging Python, PyTorch, and TensorFlow. My work includes speaker verification for security applications and voice-based user interfaces, frequently addressing noise robustness using deep learning architectures such as ECAPA-TDNN and x-vectors. I am fully comfortable working with open datasets like VoxCeleb and augmenting them with noise samples from sources such as MUSAN to simulate challenging acoustic conditions. Here’s my proposed approach: - Dataset assembly and curation from publicly available sources, applying advanced noise augmentation (additive background noise, reverberation, and SNR mixing) to simulate café, street, and industrial settings. - Model selection and training, likely starting with ECAPA-TDNN or x-vector, and experimenting with hybrid architectures for further robustness if needed. - Rigorous validation using a noisy, withheld test set to ensure ≥90% accuracy. - Delivery of a lightweight Python API/CLI tool that processes wav/mp3 files and outputs speaker ID with confidence scores. - A concise technical report covering data, augmentation, model choices, training, and validation. To ensure we hit your real-world performance goals, could you share more about your expected number of speakers and intended deployment environment? I look forward to collaborating and am confident we can create a solution that excels even in the toughest acoustic scenarios.
$140 USD in 5 days
4.6
4.6

Hi, The key challenge here is not speaker recognition in clean audio, but preserving discriminative speaker embeddings when the recording is contaminated by café noise, machinery, traffic, reverberation, and other real-world interference. I’d build the pipeline in Python/PyTorch using an ECAPA-TDNN or x-vector style encoder, then train/fine-tune with controlled noise augmentation across multiple SNR levels. Publicly licensable speech/noise datasets would be used, with all sources documented. The validation setup would include a withheld noisy test set so the reported accuracy reflects the actual target environment rather than clean benchmark conditions. I’d also calibrate confidence scores so low-certainty identifications can be rejected instead of forcing a speaker ID. Deliverables will include the trained model, preprocessing/inference pipeline, CLI or lightweight API accepting WAV/MP3 input, speaker ID + confidence output, and a concise technical report covering data, augmentation, training and validation results.
$185 USD in 5 days
4.7
4.7

I’ve built speaker verification pipelines using ECAPA-TDNN and x-vector backbones on VoxCeleb1/2 and LibriSpeech subsets under noisy augmentations—so this is right in my wheelhouse. I’ll train a ResNet-34 backbone with additive noise and reverberation from MUSAN and RIR datasets, then fine-tune with SNR dropouts down to 0 dB. Validation will use a held-out subset of VoxCeleb1-E with cafés and street noise mixed at random SNRs. The model exports to TorchScript for a Flask CLI that takes wav/mp3, resamples to 16 kHz, and returns speaker ID plus cosine similarity score. Questions: do you have a target speaker list or should I assume open-set identification? I can start immediately. Thanks, Andrii.
$200 USD in 3 days
4.4
4.4

Hi, I am a Python and machine learning developer with 8 years of rich experience in software development, with a background in audio processing, deep learning, model evaluation, and API development. I am familiar with Python, PyTorch, TensorFlow, speaker embeddings, ECAPA-TDNN, x-vector approaches, audio preprocessing, noise augmentation, and confidence scoring. I understand the goal is speaker identification only, with reliability in difficult real-world noise rather than transcription. I would build the pipeline around speaker embeddings, aggressive noise and SNR augmentation, and evaluation against a separate noisy test set, then expose the final model through a lightweight CLI or API returning speaker ID and confidence. I'm an individual freelancer and can work on any time zone you want. Please contact me with the best time for you to have a quick chat. Thanks. Emile.
$250 USD in 7 days
4.5
4.5

Nice to meet you , My name is Anthony Muñoz, I express my interest in working on your project after carefully reading the requirements and concluding that they match my area of knowledge and skills. I am currently the lead engineer for the IT agency DSPro and I have more than 10 years of experience in the field. I have successfully completed a large number of similar jobs and I consider your project to be a challenge in which I would like to work and be able to make it a reality. Please feel free to contact me, it will be my pleasure to help you. I greatly appreciate the time provided and I remain attentive to any questions or concerns. Greetings
$134 USD in 7 days
4.6
4.6

Hi, I can build a Python based speaker recognition pipeline using PyTorch with a proven architecture such as ECAPA TDNN, including noise augmentation and preprocessing for challenging environments. The solution will include a trained model, lightweight API or CLI returning speaker ID and confidence, plus validation on a withheld noisy test set targeting 90 percent or higher accuracy. I can also provide the dataset details, training approach, technical report, and a clear timeline before starting. Best regards, Huzaifa
$120 USD in 5 days
4.6
4.6

I can develop a reliable speaker recognition model capable of identifying human voices in clatter-rich environments. Using PyTorch and deep learning models like ECAPA-TDNN, the solution will focus on handling complex noise scenarios. We'll use datasets such as VoxCeleb2 and LibriSpeech, both public and suitable for this task. The training will incorporate SpecAugment and RIR & Noise Augmentation techniques to prepare the model for real-world noise conditions. This ensures the model is effective even when signal-to-noise ratios are low. The project is scheduled over five weeks, starting with data collection and preprocessing in the first week. I'll develop the model and integrate noise augmentation in the second week. The third week is dedicated to validating the model with a noisy test set, targeting at least 90% accuracy. In week four, I'll build an API/CLI demo for easy deployment, wrapping up with a comprehensive technical report in the final week. The cost will range from $15,000 to $25,000. If this approach aligns with your needs, let’s discuss the next steps. Shawana Younus
$140 USD in 7 days
4.0
4.0

The background clatter you're describing—café, factory, street—is exactly where embedding-based systems like x-vector tend to degrade fastest, since real-world SNR drift doesn't match clean augmented training data unless the noise mixing is done carefully at multiple SNR levels rather than a single fixed ratio. My plan: build the pipeline in Python around an ECAPA-TDNN backbone (via SpeechBrain or a PyTorch reimplementation), train on VoxCeleb1/2 for base speaker embeddings, then apply on-the-fly noise augmentation using MUSAN and a subset of DEMAND for the café/street/factory-type profiles you mentioned, mixed at randomized SNR ranges rather than a single level. Evaluation would report accuracy specifically on a held-out noisy split at varying SNR bands so the 90% target is verifiable rather than a single blended number. Delivery would be a FastAPI endpoint plus a CLI, both taking a wav/mp3 and returning speaker ID and confidence. One question: is this closed-set identification against a fixed roster of known speakers, or open-set verification where unseen speakers must be rejected—the model architecture and threshold logic differ meaningfully between the two. Happy to start on the data pipeline and augmentation setup as soon as you confirm which case applies.
$30 USD in 4 days
3.7
3.7

Dear Client, I’m an experienced full-stack developer with over 10 years of experience in web and mobile application development, specializing in building scalable, responsive, and high-performance solutions for diverse business needs. I understand you are looking for a reliable developer to build or improve your project, including web or mobile applications similar to CRM, dashboards, or APIs, and I have worked on similar solutions successfully. My skills in React, Vue, Laravel, PHP, Python, REST APIs, and database design ensure efficient and high-quality delivery. Feel free to share more details or ask questions. I’m ready to refine my approach to match your exact requirements. Looking forward to working with you. Best regards, Md Ruhul Ajom
$100 USD in 3 days
5.2
5.2

With respect to your project, I understand and appreciate the need for a robust solution that can accurately identify human voices in a wide range of noisy environments. My expertise in machine learning combined with Python fluency renders me highly equipped to deliver an optimized model. For this task, I would recommend building a custom acoustic model using recurrent neural networks (RNN), which is particularly effective in noise endurance. This model will leverage TensorFlow framework for seamless integration. My experience extends to working with publicly licensable datasets, and I adhere strictly to all licensing protocols. Moreover, deploying relevant data augmentation techniques, such as SpecAugment and Wav2Vec 2.0, will significantly improve the model's ability to handle noises effectively. After training the model on a diverse set of noisy audio datasets with a strong emphasis on robustness, I will provide you with not just a trained speaker recognition model capable of handling noisy audio but a lightweight API or CLI demo that accepts wwav/mp3 files and returns the speaker ID along with a confidence score. Finally, my commitment to timely delivery and ongoing support make me the right fit for this project
$180 USD in 4 days
3.9
3.9

Hello, I read your project description carefully, and I’m very glad to see that it aligns closely with my background and experience. I have a solid foundation in machine learning, with particular experience and interest in sequential data processing, signal processing, and deep learning. With my background in DSP and mathematics, I’m comfortable working with raw audio signals, extracting meaningful features, and designing advanced end-to-end neural network architectures for speech and audio-related applications. I’m experienced with both foundational data-science libraries such as NumPy, Pandas, SciPy, and Scikit-learn, as well as deep learning frameworks including TensorFlow, Keras, and PyTorch. This allows me to work across the entire pipeline, from raw data processing and feature engineering to model development, training, evaluation, and optimization. I also place strong emphasis on code quality and maintainability. I can provide a well-structured pipeline, clear documentation, readable and optimized code, and helpful comments to make the project easier to understand and extend. I would be happy to learn more about your requirements and discuss how my background could contribute to your project. I look forward to hearing from you. Thank you for your consideration.
$160 USD in 6 days
3.7
3.7

Hi! there - Abror here "NOISY SPEAKER RECOGNITION" — you need reliable speaker ID when real-world noise makes clean voice embeddings unstable. I’d start with an ECAPA-TDNN style embedding model and train with controlled noise augmentation using publicly licensable speech and noise datasets. I’d keep the noisy test set completely separate from training, so the 90% target reflects actual performance rather than memorized conditions. The demo can stay lightweight: WAV/MP3 in, speaker ID and confidence score out, with the report showing preprocessing, training setup and results across different noise levels. How many known speakers will the final model need to identify, and do you already have clean recordings for those speakers?
$30 USD in 1 day
3.0
3.0

Kakanj, Bosnia and Herzegovina
Payment method verified
Member since Aug 3, 2025
$30-250 USD
$10-30 USD
$30-250 USD
$30-60 USD
$10-30 USD
$8-15 USD / hour
€30-250 EUR
$250-750 USD
£250-750 GBP
$250-750 USD
£250-750 GBP
$30-250 USD
$500 USD
$30-250 USD
€30-250 EUR
$10-30 USD
$2-8 USD / hour
€8-50 EUR
₹600-1500 INR
$30-250 USD
$15-25 USD / hour
$30-250 USD
min $50 USD / hour
₹1800-2800 INR / hour
₹12500-37500 INR