
Closed
Posted
Site Reliability Engineer (Contract) Location: India, remoteHours: 09:00–18:00 UK time, Mon–Fri (13:30–22:30 IST summer, 14:30–23:30 IST winter)Experience: 5–6 yearsCloud: Azure and AWS We run production infrastructure for UK clients during UK office hours, plus out-of-hours support and fixes for US clients. The platform team is two engineers. This is the reliability seat: you own how we know the platform is healthy, how we find out when it isn't, and how fast we recover. Your counterpart owns delivery pipelines, and you cover each other. We mean SRE in the substantive sense, writing code to remove operational work rather than only operating things. If that isn't how you want to spend your week, the DevOps seat is the better fit. Technical Cloud. Strong depth in either Azure or AWS, working knowledge of the other. We don't expect equal depth in both. AWS: EC2, S3, IAM, Lambda, CloudWatch, VPC, RDS, EKS/ECS Azure: VMs, Blob Storage, Entra ID/RBAC, Functions, Azure Monitor, VNet, Azure SQL, AKS Observability. Design and run logging, metrics, tracing and alerting. Prometheus, Grafana, ELK/OpenSearch, Azure Monitor, Datadog or similar. Dashboards that work during an incident, not just in a review. Software engineering. Python or Go at production standard: tested, reviewed, maintained by others. Scripting alone isn't enough for this seat. Kubernetes. Production experience debugging workloads, not only deploying them. Reliability practice. Defining SLIs and SLOs, capacity forecasting, performance investigation across the application and database boundary, cloud cost efficiency. Foundations. Strong Linux administration and deep troubleshooting. Solid Terraform. Networking: DNS, TCP/IP, HTTPS, load balancers, security groups, TLS. Enough CI/CD knowledge to cover the other seat. Useful to have: OpenTelemetry and distributed tracing; formal SLO or error-budget practice; FinOps; database performance tuning; GitOps with ArgoCD or Flux; AZ-305 or AWS Solutions Architect. Procedural You lead incidents during UK hours and on your cover weeks, then run the post-incident review. Blameless, with root cause identified and follow-up actions tracked to completion. Alerting is tuned continuously. Every page should be actionable, and noisy alerts are treated as defects. A rota this small can't absorb false pages. Production readiness review before new services or clients go live. Runbooks are deliverables. Whoever picks up a page overnight should be able to work through it without calling the other engineer. Recurring manual work gets identified and removed with code rather than absorbed. Infrastructure changes go through code review and pipelines, alongside the DevOps engineer. The working day Fixed hours in UK time, so the window follows UK clock changes rather than drifting. Full overlap with your counterpart and with the UK team, so most work is collaborative rather than solo. Out-of-hours cover for US clients on alternating weeks, included in the scope of the engagement. Response expectations are set by severity, and we track page volume with the intent of keeping it low. A real objective of this role is making that rota quieter: better alerting and automated remediation so overnight breakage is caught and where possible fixed without a human. Agile/Scrum cadence with a distributed team: standups, planning and retros in UK hours. Written English at CEFR C1 or equivalent (IELTS 7.0+, or comparable professional experience). A large share of communication is asynchronous, and you'll regularly hand a live incident to someone who was asleep when it started. BYOD. You supply your machine; all production access runs through a managed virtual desktop, with nothing sensitive stored locally. Disk encryption, supported OS, screen lock and endpoint protection required, plus reliable broadband and a backup connection for on-call weeks.
Project ID: 40676135
51 proposals
Remote project
Active 3 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
51 freelancers are bidding on average $12 USD/hour for this job

Nice to talk you , After reading in detail the requirements of your project and concluding that they match my areas of knowledge and skills, I would like to introduce myself. My name is Anthony Muñoz and I am the lead engineer for DS Pro IT agency. I have worked for over 10 years in Backend and software development and have successfully done multiple jobs. It will be a pleasure to work together to make your project a reality. Please feel free to contact me. I´m looking forward to working with you. I really appreciate your time and remain attentive to any request or question. Greetings
$12 USD in 40 days
5.7
5.7

Hi, Your biggest SRE challenges are keeping production healthy, reducing noisy alerts, recovering quickly from incidents, and eliminating recurring manual operational work through automation. I can help build a reliability-focused environment where monitoring is actionable, incidents are easier to troubleshoot, and the platform becomes progressively more stable. I have 16+ years of experience in AWS, Azure, Linux, DevOps, SRE, Kubernetes, Terraform, CI/CD, security, and production infrastructure. My hands-on expertise includes EC2, EKS/ECS, RDS, CloudWatch, Azure VMs, AKS, Azure Monitor, Prometheus, Grafana, ELK, Docker, networking, and infrastructure automation. I can take ownership of observability, SLIs/SLOs, alert tuning, incident response, RCA, capacity and performance analysis, Kubernetes troubleshooting, cloud cost optimization, and production readiness. I also use Python/Bash automation and Terraform to remove repetitive operational work and improve reliability. I’m comfortable working UK hours, collaborating with distributed teams, participating in incident rotations, and maintaining clear runbooks and documentation. I’m available to start immediately and can discuss the role through the Freelancer call option. Best regards, SaD
$15 USD in 40 days
5.3
5.3

With a strong foundation in backend development, DevOps engineering and over 5 years of motivation - Your team deserves the competence to keep its platform up and thriving. My broad experience across various frameworks (e.g., AWS, Terraform, Docker, Kubernetes) allows me to easily navigate cloud operations, pipeline setups, and configuration management for peak productivity, a fundamental part of delivering a sustainable and reliable platform. My deep expertise with Amazon Web Services including EC2, S3, IAM, Lambda and cloud monitoring tools such as Prometheus and Grafana gives me the necessary strength in cloud infrastructure for your project. Above this benefit that my cloud knowledge brings is also my extensive experience with Linux administration which will prove quite valuable in addressing troubleshooting needs that may arise during your UK office hours. In conclusion my name is Ashish, an AWS Certified professional with the agility skills you are looking for; adaptability to Drifted UK clock changes or; fixed hours in UK time and enough CI/CD knowledge to cover the other seat. I'm ready to be the key to your solution: dependable through dark incidents, effective at brainstorming post-incident courses of actions, diligent about automation and prioritizing customer satisfaction with productive communication at stand ups; planning or retros,. Choose me to make your rota quieter so that it remains unquestionably effective irrespective daytime or half awaken situations.
$15 USD in 40 days
5.5
5.5

Hi, I’m a Site Reliability/DevOps Engineer with 5+ years of experience managing production infrastructure across Azure and AWS. I have strong hands-on expertise in Kubernetes, Terraform, Linux, networking, CI/CD, monitoring, logging, and incident troubleshooting. I build actionable observability using Prometheus, Grafana, ELK, and cloud-native monitoring, and I’m comfortable defining SLI/SLOs, improving alert quality, and automating repetitive operational work with Python/Bash. I’ve handled production incidents, root-cause analysis, performance issues, and reliability improvements in distributed environments. I focus on engineering solutions that reduce manual operations and make systems more resilient, observable, and cost-efficient. I’m available for UK-hour overlap and on-call rotations.
$12 USD in 40 days
5.3
5.3

Hi, I have 6+ years of experience as a DevOps/SRE engineer with strong depth in Azure and AWS (EC2, S3, IAM, Lambda, CloudWatch, VPC, RDS, EKS/ECS), Kubernetes debugging, Prometheus/Grafana/ELK, Terraform, and production Python. I design SLIs/SLOs, tune actionable alerting, lead blameless post-incident reviews, and write code to eliminate toil rather than just operate systems. My approach will be: 1. Own platform health: design and run observability (metrics, logs, tracing, alerting) so dashboards and pages stay actionable during incidents. 2. Drive reliability practice: define SLIs/SLOs, capacity forecasting, performance investigation across app/DB boundaries, and cloud cost efficiency. 3. Lead UK-hours incidents and cover weeks, run post-mortems with tracked follow-ups, and continuously remove noisy alerts and recurring manual work via automation. 4. Partner on production readiness reviews, maintain runbooks, and ensure infrastructure changes go through code review and pipelines. Deliverables: quieter on-call rota through better alerting and automated remediation, production-ready observability and runbooks, and seamless cover with the delivery counterpart. Quick questions: 1. Which cloud currently carries the larger production footprint for UK vs US clients? 2. What is the current page volume and primary source of overnight noise? You’ll end up with a more reliable platform and a quieter rota that scales with the team. Best Regards, Rahul
$10 USD in 40 days
4.0
4.0

Hi, I’m an experienced SRE/DevOps engineer with strong hands-on expertise in AWS/Azure, Linux, Kubernetes, Docker, Terraform, CI/CD, and production troubleshooting. I can build reliable observability with Prometheus/Grafana/CloudWatch/Azure Monitor, improve alerting, automate recurring operational work using Python, and lead incident response/postmortems. I’m comfortable with UK-hour overlap and out-of-hours rotation, and I focus on reducing incidents and making systems more reliable, secure, and cost-efficient. I’d be happy to discuss how I can contribute to your platform team.
$10 USD in 40 days
3.5
3.5

Hi , I'm having 14 + years of experience in multi cloud environment. Having experience to manage 20k vm across the globe. This is my day-to-day job. Let me know if i can help you
$12 USD in 15 days
3.3
3.3

SolutionzHere has placed senior SREs for UK/US clients with Azure + AWS stacks. Our engineers own observability (Prometheus/Grafana/ELK/Datadog), write production-grade Python/Go, debug K8s workloads, define SLIs/SLOs, and drive blameless incident reviews with actionable runbooks. Your budget is below market for 5–6 years’ SRE in UK hours; typical is USD 22–30/hr. We propose USD 26/hr, full-time (09:00–18:00 UK), with alternating US on-call cover and a 30-day trial. Question: What’s your current alert volume per week and which observability stack (Prometheus/Grafana, Datadog, Azure Monitor) is primary today?
$25 USD in 40 days
3.3
3.3

This role is clearly focused on reliability engineering rather than general DevOps work: strong observability, incident response, production debugging, SLO ownership, and reducing repetitive operational work through code. I would approach the environment by first understanding the current AWS/Azure architecture, Kubernetes workloads, alerting stack, Terraform structure, incident history, and existing operational pain points. From there, I would tighten SLIs/SLOs, improve dashboards and alert quality, identify noisy or non-actionable pages, and automate recurring recovery tasks using Python or Go. For production incidents, I am comfortable working systematically across Linux, networking, containers, Kubernetes, cloud services, application logs, and database behavior. I would also keep runbooks and post-incident actions practical so another engineer can take over an issue without needing verbal context. Infrastructure and remediation changes would stay code-reviewed and reproducible through Terraform and CI/CD rather than manual console changes. I am also comfortable with full UK-hours overlap and a collaborative two-engineer platform setup. Which cloud currently carries most of the production workload, and what observability stack and Kubernetes platform are already in place?
$9 USD in 40 days
3.1
3.1

Hi, I’m an experienced Linux/Cloud DevOps engineer with strong hands-on experience in AWS, Azure, Kubernetes, Terraform, Linux administration, CI/CD, monitoring, networking, and production troubleshooting. The reliability-focused nature of this role particularly interests me. I can take ownership of observability, alerting, incident troubleshooting, performance issues, infrastructure reliability, and operational automation, while working closely with the delivery-focused engineer. I’m comfortable working with EC2, VPC, IAM, RDS, S3, ECS/EKS, Azure VMs, VNet, RBAC, AKS, CloudWatch, Azure Monitor, Prometheus/Grafana, and production Kubernetes environments. I also focus on reducing recurring operational work through automation and maintaining clear runbooks and post-incident documentation. I’m comfortable with the required UK working hours and remote collaboration, including structured incident handovers and on-call responsibilities. I’m available to start immediately and would be happy to discuss your current infrastructure, reliability goals, and on-call expectations
$15 USD in 40 days
3.0
3.0

Greetings, I’m an **AWS Certified DevOps Engineer – Professional** with 6+ years of hands-on DevOps/SRE experience supporting production environments where reliability, observability, incident response, and automation are critical. My strongest cloud expertise is AWS, including **EC2, EKS/ECS, RDS, VPC, IAM, S3, Lambda, and CloudWatch**, with working experience across Azure. I also have strong experience with **Linux, Kubernetes, Terraform, Python/Bash, Prometheus, Grafana, ELK/OpenSearch, networking, CI/CD, and production troubleshooting**. My SRE approach includes: * Building actionable monitoring, dashboards, and alerts * Defining SLIs/SLOs and tracking service health * Debugging Kubernetes, application, and database performance issues * Leading incident response, RCA, and post-incident improvements * Automating recurring operational work to reduce toil * Creating practical runbooks for independent incident handling * Managing infrastructure through Terraform and CI/CD I’m comfortable taking ownership during incidents, working in a small platform team, and providing clear technical handovers. I can work **09:00–18:00 UK time from India** and participate in the alternating out-of-hours support rotation. I’d be happy to discuss your current SRE challenges and reliability goals. Best regards, Ankush
$10 USD in 40 days
2.7
2.7

SRE work shouldn’t mean firefighting—it should mean building systems that fight fires themselves. You need someone who treats operational toil as technical debt, not a job description. Your focus on actionable alerts, runbooks, and eliminating manual work aligns with how I’ve operated in production environments for global SaaS platforms. At Skybin, I architected AWS/Azure cloud infrastructure with automated monitoring, recovery workflows, and strict CI/CD enforcement—cutting incident volume by making systems self-healing. I’ve used Docker, Kubernetes, and Terraform to standardize deployments and reduce configuration drift. I’d start by auditing current alerting and runbooks, identifying top toil sources, and implementing automated remediations using existing pipelines. Observability would be tied to real user impact via SLIs/SLOs. Happy to review your current setup and outline a first-step automation plan.
$12 USD in 40 days
2.4
2.4

Hi, We have a senior Site Reliability Engineer available and ready to start, with 5+ years of hands-on experience in cloud infrastructure, automation, monitoring, CI/CD and production support. The resource can work during your required time zone and coordinate directly with your team. Questions: 1. Which cloud platform is currently being used — AWS, Azure, or GCP? 2. What is your current deployment setup — Kubernetes, Docker, VM-based, or serverless? 3. Which areas are the immediate priority: reliability, monitoring, cost optimization, or CI/CD? Matching Skills & Tools: Experience: 5+ years SRE / DevOps Cloud: AWS, Azure, GCP Containers: Docker, Kubernetes, Helm CI/CD: Jenkins, GitHub Actions, GitLab CI/CD IaC: Terraform, Ansible Monitoring: Prometheus, Grafana, CloudWatch Logging: ELK, Loki, Cloud Logs Scripting: Python, Bash, Shell Databases: PostgreSQL, MySQL, MongoDB Version Control: Git, GitHub, GitLab Security: IAM, Secrets Management, Access Control Other: Incident Management, Automation, Performance Optimization Technical approach: • Review infrastructure and production environment. • Improve monitoring, alerting, reliability and automation. • Optimize deployments, scaling and resource usage. • Strengthen CI/CD, security and disaster recovery. • Document infrastructure and operational procedures. Our senior resource is ready to start, with additional Cloud, Backend, QA and DevOps support available when required. Best Regards, Nikhilesh Team Qloron
$15 USD in 40 days
2.1
2.1

Hi — Diego Vinicius here from Brazil. I’m an experienced Fullstack/DevOps engineer with hands-on experience across AWS, Linux, Docker, Kubernetes, Terraform, CI/CD, networking, and production troubleshooting. I’m comfortable treating SRE as engineering: automating operational work, improving observability, and reducing recurring incidents rather than simply maintaining infrastructure. I can work with AWS/Azure environments, Kubernetes workloads, monitoring/logging, incident investigation, database performance, infrastructure-as-code, and reliability improvements. Q1 – Which cloud currently hosts the majority of your production workloads, AWS or Azure? Q2 – What observability stack is currently in place? Q3 – How is the out-of-hours on-call rotation structured? I’m comfortable with UK working hours, collaborative Agile workflows, incident ownership, documentation/runbooks, and long-term reliability improvements.
$20 USD in 40 days
1.7
1.7

Hello, I’ve carefully reviewed your requirements and have the expertise to deliver this project with high quality, on time, and to your expectations. With 6+ years of hands-on experience in Python automation, social media growth, and AI-driven workflows, I’m confident I can deliver the results you need. With over six years in cloud operations, I specialize in building resilient infrastructures on AWS and Azure. I’ll begin by auditing current monitoring, defining SLIs/SLOs, and tightening alert rules so every page is actionable. Next, I’ll automate incident runbooks in Python, integrate Prometheus/Grafana and Datadog, and set up automated remediation scripts for common failure patterns. I’ll review capacity and cost, apply Terraform best‑practice, and conduct a production readiness run‑through for new services. During UK hours I’ll lead on‑call incidents, run post‑incident reviews, and drive continuous improvement. My background in Python, Kubernetes, and cloud‑native observability tools aligns with your needs. I will also set up automated cost monitoring dashboards and implement FinOps practices to keep spend in line with performance goals. Additionally, I will document runbooks and conduct training sessions for the team to ensure smooth handovers. Let’s discuss how I can help reduce the on‑call load and improve uptime. Looking forward to discussing the project details further on chat. Best regards, NAVEEN THAKUR
$8 USD in 40 days
0.8
0.8

As a seasoned Full Stack Developer with over 5 years of experience, I have a deep understanding of cloud infrastructure and platform reliability that aligns perfectly with your SRE requirements. My proficiency with both Azure and AWS ensures that regardless of your cloud platform preference, I can effectively design, deploy, and monitor your system for optimal performance and health. Additionally, my expertise with key technologies like EC2, S3, IAM, Lambda, CloudWatch on AWS as well as VMs, Blob Storage, Entr ID/RBAC etc., on Azure make me versatile enough to handle complex tasks across different environments. One of my strengths lies in my ability to write production-standard code in Python or Go, ensuring that it is tested, reviewed, and easily maintainable by other team members. I've extensively used logging systems like Prometheus, Grafana, ELK/OpenSearch along with various monitoring tools including Azure Monitor and Datadog. These nuances are critical in handling incidents effectively, reducing false alerts noise and delivering outcomes based on thorough post-incident follow-up reviews. I also possess strong skills in Kubernetes for debugging workloads allowing quick restoration alongside other relevant competencies such as capacity forecasting cloud cost efficiency. let's connect to discuss how I can contribute to the improvement of your infrastructure environment!
$12 USD in 40 days
1.4
1.4

Static alert thresholds in CloudWatch and Azure Monitor can miss gradual degradation that later triggers noisy pages. I’ll replace those with dynamic baselines built from Prometheus metrics and feed them into Grafana alerts that auto‑adjust to traffic patterns. Then I’ll add a small Python script that writes remediation steps to a Terraform module, turning recurring manual fixes into version‑controlled code. One common mistake is treating alerts as a checklist instead of a signal, which leads to alert fatigue. You’ll see fewer false pages, faster root‑cause isolation, and runbooks that let anyone handle a night‑time incident without calling the other engineer.
$12 USD in 40 days
0.0
0.0

Hello, This role requires true SRE ownership: measurable reliability, actionable observability, automated remediation, and engineering that removes recurring operational work. I would first map the AWS/Azure infrastructure, Kubernetes workloads, dependencies, current incidents, alert noise, and existing SLOs/runbooks. Then I’d strengthen monitoring with meaningful metrics, logs, tracing, and incident-focused dashboards using tools such as Prometheus, Grafana, CloudWatch, Azure Monitor, or OpenTelemetry. For reliability issues, I’d troubleshoot across Linux, Kubernetes, networking, applications, and databases, using production-grade Python/Go automation to eliminate repetitive manual work. Terraform and Git-based reviews would keep infrastructure changes controlled and reproducible. Incident response would include clear severity handling, root-cause analysis, blameless postmortems, and tracked remediation. Production-readiness reviews and practical runbooks would also make out-of-hours support safer and faster. The objective would be measurable: fewer false alerts, faster recovery, better capacity planning, improved cloud efficiency, and progressively quieter on-call weeks. Which cloud currently carries the larger production workload, AWS or Azure, and what is the biggest reliability issue you want addressed first?
$12 USD in 40 days
0.0
0.0

Hi, I’m interested in the Site Reliability Engineer role. I have strong experience with cloud infrastructure, Linux, Kubernetes, Terraform, monitoring, troubleshooting, and CI/CD. I focus on improving reliability, reducing manual work, and making alerts and incidents easier to manage. I’m comfortable working with distributed teams, handling production issues, and following UK working hours. I’m also committed to clear communication, proper documentation, and reliable delivery. I’d be happy to discuss your infrastructure and how I can support your platform team. Thank you!
$12 USD in 40 days
0.0
0.0

I’m a Senior DevOps/SRE engineer with 10+ years of experience running production AWS and Azure infrastructure, with strong hands-on skills in Kubernetes, Terraform, Linux, Python, networking, observability, and incident response. I’ve managed production environments using EKS/ECS, AKS, RDS/Azure Database, CloudWatch, Azure Monitor, Prometheus, Grafana, OpenSearch, and Datadog. I focus on reducing operational work through automation, improving alert quality, and making incidents easier to diagnose and recover from. I’m comfortable owning incidents, writing post-incident reviews, creating runbooks, defining SLI/SLOs, troubleshooting application/database performance, and working closely with a delivery-focused DevOps engineer. I can work **09:00–18:00 UK time, Monday–Friday**, including the required alternating out-of-hours coverage. I’m based in India and can provide full UK-hour overlap. I’d be happy to discuss how I can help make the platform more reliable while reducing manual operational work and unnecessary pages. Thanks
$15 USD in 40 days
0.0
0.0

Pune, India
Member since Apr 19, 2012
$8-10 USD / hour
₹12500-37500 INR
₹12500-37500 INR
₹12500-37500 INR
$250-750 USD
$8-15 USD / hour
₹12500-37500 INR
$100-200 USD
$30-250 USD
₹12500-37500 INR
£20-250 GBP
₹750-1250 INR / hour
$3000-5000 AUD
₹12500-37500 INR
$250-750 USD
$30-250 USD
₹37500-75000 INR
₹75000-150000 INR
$30-250 CAD
$250-750 USD
₹1500-12500 INR