
Closed
Posted
Paid on delivery
Infra Engineer – SRE (Kubernetes) (INDIAN EMPLOYEE ONLY) About the Role We are seeking a skilled Site Reliability Engineer specializing in Kubernetes to join a Global Infrastructure team. This role is hands-on and critical to ensuring the stability, efficiency, and reliability of large-scale high-performance AI/ML clusters in data centers. The ideal candidate will bring expertise in system-level troubleshooting, AI cluster maintenance, and operational excellence to ensure maximum performance for infrastructure environments. Experience with large-scale infrastructure automation is considered a strong plus. Responsibilities * Design, implement, and maintain scalable AI/ML infrastructure solutions. * Proactively monitor GPU cluster health, performance, and troubleshoot issues across compute, accelerator, networking, and storage systems. * Automate deployment, configuration, and management of infrastructure resources. * Manage GPU node lifecycle workflows, including provisioning, scaling, maintenance, decommissioning, and upgrades. * Implement CI/CD pipelines for infrastructure deployment and orchestration. * Ensure security, compliance, and operational best practices across infrastructure environments. * Manage incident response related to infrastructure resources, including GPU, CPU, storage, and network components. * Handle customer provisioning requests for GPU resources, including onboarding, configuration, and troubleshooting. * Resolve customer service requests related to infrastructure and platform operations while maintaining high customer satisfaction. * Stay current with emerging GPU hardware and software technologies and integrate improvements where appropriate. * Support regional and international travel requirements to data center locations when necessary. Qualifications * Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field. * 3+ years of experience in data center operations, infrastructure engineering, systems engineering, or site reliability engineering. * Proven experience with infrastructure automation tools such as Terraform and Ansible. * Strong experience with Kubernetes and container orchestration technologies. * Familiarity with NVIDIA GPU Operator, NVIDIA Network Operator, CNI, CSI, and similar Kubernetes ecosystem tools. * Experience with job scheduling systems such as Slurm. * Strong Linux system administration skills. * Proficiency in scripting and automation using Python and Bash. * Experience with observability and monitoring platforms such as Prometheus, Grafana, and Loki. * Knowledge of GPU architectures, NVIDIA CUDA, NCCL, and AI/ML infrastructure is a strong advantage. * Strong troubleshooting and root-cause analysis skills with the ability to analyze logs, metrics, and system performance data. * Excellent communication, collaboration, and problem-solving abilities. Preferred Skills * Large-scale Kubernetes cluster operations. * AI/ML infrastructure and GPU cluster management. * Infrastructure-as-Code (IaC) and automation-first mindset. * Production incident management and reliability engineering. * Data center operations and hardware troubleshooting. * CI/CD platform design and implementation. Meeting every qualification is not required. Candidates with strong technical foundations, relevant experience, and a passion for building reliable large-scale infrastructure are encouraged to apply.
Project ID: 40535266
19 proposals
Remote project
Active 1 day ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
19 freelancers are bidding on average ₹25,639 INR for this job

The difficult part in this role isn't Kubernetes itself, it's keeping GPU-heavy clusters stable when issues span multiple layers at once: Kubernetes, networking, storage, drivers, CUDA/NCCL, and hardware. I've spent a lot of time troubleshooting Linux and container infrastructure where the root cause wasn't obvious from cluster metrics alone. My approach would be automation-first from day one: Terraform and Ansible for repeatable infrastructure changes, strong observability with Prometheus/Grafana/Loki, and clear operational workflows around provisioning, upgrades, and incident response. For AI workloads, GPU lifecycle management and monitoring are usually where reliability problems surface first, so I'd focus heavily there. Are your GPU clusters currently running on self-managed Kubernetes or a managed platform? That affects how I'd structure node lifecycle automation and upgrade strategy.
₹12,500 INR in 7 days
3.6
3.6

I am excited to work on your SRE for AI/ML Infrastructure project. With strong experience in Kubernetes and automation, I can help ensure the reliability and performance of your AI/ML clusters. My skills in Python and Linux allow me to automate deployment and configuration effectively. I have a proven track record with infrastructure automation tools like Terraform, ensuring streamlined management of resources. To maintain optimal GPU cluster health and troubleshoot issues, I will implement CI/CD pipelines aligned with operational best practices. How do you currently monitor and maintain GPU cluster performance? What specific challenges do you face in your AI/ML infrastructure? Do you have any existing automation tools you prefer to use? Looking forward to potentially collaborating on this critical project.
₹12,500 INR in 13 days
3.5
3.5

Hi, I'm a full-stack engineer with hands-on experience running Kubernetes for AI/ML and GPU workloads, plus large-scale infrastructure automation. My approach: - Monitor GPU cluster health (compute, accelerator, networking, storage) with Prometheus/Grafana + custom alerting to catch issues proactively - Automate node lifecycle: provisioning, scaling, maintenance, decommissioning and upgrades via IaC (Terraform/Ansible) and GitOps - Harden K8s scheduling for high-performance AI clusters (node affinity, taints/tolerations, GPU device plugins) - Build self-healing runbooks and automated remediation for common cluster failures - Drive operational excellence: SLOs, capacity planning, and post-incident reviews I'm India-based and available hands-on. Quick question: how many GPU nodes are in the cluster, and is it on-prem data center, cloud (EKS/GKE), or hybrid?
₹25,000 INR in 20 days
3.7
3.7

Hi, Krishna here from Delhi. As an experienced member of a team that specializes in delivering artificial intelligence solutions and investing in reliable infrastructure, I can offer immense value to your project. Though I am yet to have hands-on professional experience with integrations with NVIDIA infrastructure and SLURM, my solid understanding in toolchains such as Kubernetes and Ansible align well with the requirements of this role. Furthermore, my adaptability and commitment to apprising myself of novel technologies will ensure that any existing gaps are promptly closed, resulting in efficient management and deployment of the system. My expertise extends beyond just competence in languages such as Python and Bash; it also encompasses a deep comprehension of AI/ML clusters maintenance and optimization. Precisely, this includes procedural skills like implementing both CI/CD pipelines for seamless orchestration as well as using observability platforms like Grafana to monitor performance & build logs for root-cause analysis - all tertiary virtues in the effectuating operational excellence for large-scale systems like those you have.
₹25,000 INR in 7 days
1.7
1.7

I am interested in the position. I have 3+ years experience as a Devops engineer with extensive experience in Kubernetes cluster creation and management. Apart from the AI and GPU management I have experience in all other skills that you listed. Lets connect.
₹70,000 INR in 30 days
1.4
1.4

Hi, there. I’m Susie Kalson, a Site Reliability Engineer with over 7 years of experience in infrastructure management, particularly in Kubernetes and AI/ML environments. I understand that maintaining stability and performance in large-scale GPU clusters is critical for your operations, and you need a reliable solution to streamline these processes. ✅ Automate deployment and management of Kubernetes infrastructure using Terraform and Ansible. ✅ Monitor GPU cluster health and troubleshoot issues across all systems. ✅ Manage GPU node lifecycle workflows to ensure optimal resource utilization. ✅ Implement CI/CD pipelines for efficient infrastructure deployment. ✅ Provide incident response and resolve technical issues to enhance customer satisfaction. I leverage tools like Prometheus and Grafana for observability and have a solid foundation in Linux and scripting. My familiarity with NVIDIA technologies complements my operational expertise, ensuring your infrastructure runs smoothly and efficiently. I am looking forward to work with you. Best Regards, Susie Kalson
₹12,500 INR in 5 days
0.0
0.0

Hey, I think you would like to read this. One specific detail of the project is maintaining large-scale AI/ML clusters in data centers with Kubernetes expertise. I specialize in Site Reliability Engineering for Kubernetes infrastructure. My services include designing, implementing, and automating scalable solutions to ensure peak performance. What you'll get: - Scalable AI/ML infrastructure solutions - Proactive monitoring and troubleshooting of GPU clusters - Automated deployment and management of resources - CI/CD pipeline implementation - Enhanced security, compliance, and operational practices I would like to discuss more about the project. You lose nothing. If milestones are created, payment will be fully protected, and you can use this message as proof for a full refund. Kind Regards, Riyaat
₹18,750 INR in 7 days
0.0
0.0

We've just completed a similar project, helping another team enhance their AI/ML infrastructure. We've built reliable solutions for high-performance clusters, ensuring stability and efficiency. Your goal is to optimize AI/ML infrastructure in data centers. I understand the need for clean, professional, and integrated systems to support large-scale operations. Our team specializes in Kubernetes, infrastructure automation, and system-level troubleshooting. With expertise in GPU cluster maintenance and CI/CD pipelines, we ensure seamless and efficient operations. I'd love to chat about your project! The worst that could happen is you walk away with a free consultation. Regards, Keagon.
₹18,750 INR in 7 days
0.0
0.0

I am excited about the opportunity to contribute as a skilled Site Reliability Engineer specializing in Kubernetes for your high-performance AI/ML infrastructure. With over 3 years of experience in systems engineering and infrastructure operations, I am well-equipped to design, implement, and maintain scalable solutions while ensuring maximum performance and reliability. My expertise in Kubernetes, infrastructure automation using tools like Terraform and Ansible, and strong Linux system administration skills align perfectly with your requirements. I have a proven track record in troubleshooting, automation, and managing large-scale infrastructure, making me confident in my ability to meet your expectations effectively. I am eager to bring my skills to your Global Infrastructure team, ensuring seamless operations and optimized performance for your AI/ML clusters. Let's discuss how I can add value to your project and deliver exceptional results. Regards, De-Ru
₹15,000 INR in 7 days
0.0
0.0

Hi Boss I have solid experience with Linux infrastructure, automation, production troubleshooting, and high-availability systems. Your project looks very interesting. Before we proceed, I’d like to understand a few things: Are you running a self-managed Kubernetes cluster or a managed environment? What causes most incidents right now — GPU failures, networking, storage, or scheduling issues? Is your infrastructure already automated with Terraform / Ansible, or does that need to be built? I’d like to understand the current architecture first. Looking forward to your reply.
₹25,000 INR in 30 days
0.0
0.0

I understand the critical need for a skilled Site Reliability Engineer specializing in Kubernetes to ensure stability, efficiency, and reliability of large-scale AI/ML clusters. Proactively monitoring GPU cluster health and automating infrastructure resources are key to success in this role. My expertise in infrastructure automation, Kubernetes, and system-level troubleshooting align well with the requirements of the project. With a strong background in web development, full-stack development, and automation, I am confident in my ability to contribute effectively to your infrastructure team. I already have a few thoughts on how I would approach this project and would be happy to discuss them further. How do you envision integrating emerging GPU technologies into your current infrastructure setup? Regards, Nqobani
₹18,750 INR in 7 days
0.0
0.0

Hello, I am an Infrastructure and Cloud Engineer with experience in AWS Cloud, Windows/Linux Administration, Networking, Virtualization, Databases, and Automation scripting. I have hands-on exposure to Kubernetes environments, infrastructure operations, troubleshooting, and cloud infrastructure management. My background includes working with enterprise infrastructure technologies and supporting reliable, scalable systems. For this project, I can assist with Kubernetes operations, infrastructure monitoring, automation tasks, troubleshooting, and operational support while following SRE best practices. I am eager to discuss your requirements and understand your current infrastructure setup. Looking forward to hearing from you. Regards, Aditya Lokhande
₹19,000 INR in 7 days
0.0
0.0

Hello, I am a Computer Science graduate with hands-on experience in Kubernetes, OpenShift, Linux administration, infrastructure automation, and troubleshooting through my DevOps internship. During my internship, I worked on deploying and managing OpenShift clusters, configuring OpenShift Data Foundation (ODF), integrating Ceph storage, implementing MetalLB load balancing, configuring OIDC authentication, and troubleshooting cluster-related issues. I also gained experience with containerized environments, networking, storage management, and infrastructure operations. My background aligns well with the requirements of this role, particularly in Kubernetes operations, Linux systems, infrastructure reliability, and incident troubleshooting. I am a quick learner, highly motivated, and eager to contribute to large-scale AI/ML infrastructure environments. I am confident that my technical foundation and hands-on experience make me a strong candidate for this opportunity. Looking forward to discussing how I can contribute to your team. Regards, Ajishwa Kakkar
₹25,000 INR in 10 days
0.0
0.0

I will create a pipeline using Jenkins that can do your job pretty easy. Also I will try to use Terraform for your IaC purpose
₹35,000 INR in 10 days
0.0
0.0

Hello, I am an Infrastructure Engineer with hands-on experience in Kubernetes, Linux administration, Terraform, Ansible, CI/CD, and infrastructure automation. I have worked on deploying, maintaining, and troubleshooting production Kubernetes environments while ensuring high availability, reliability, and performance. For this project, I can help with: Kubernetes cluster deployment, operations, and troubleshooting Infrastructure automation using Terraform and Ansible Linux system administration and performance tuning CI/CD pipeline implementation and optimization Monitoring using Prometheus, Grafana, and Loki Incident response, root cause analysis, and infrastructure reliability GPU node provisioning and cluster operations (where applicable) I follow Infrastructure-as-Code best practices, automate repetitive operational tasks, and focus on building stable, scalable, and secure infrastructure. I am comfortable collaborating with distributed teams and providing clear technical documentation throughout the engagement. I am available to start immediately and can provide regular progress updates to ensure timely delivery. Looking forward to discussing the project in more detail. Thank you.
₹28,500 INR in 10 days
0.0
0.0

Hi there, I am an India-based AI Infrastructure Engineer and Ph.D. Researcher specializing in high-performance computing, GPU orchestration, and Kubernetes. My daily work involves designing deterministic AI architectures and managing heavy computational bottlenecks for deep learning pipelines. Why I am the right SRE for this role: GPU Orchestration: I don't just deploy standard web apps on K8s; I manage compute-heavy workloads. I am deeply familiar with integrating the NVIDIA GPU Operator to expose underlying hardware (CUDA, NCCL) to containerized environments, ensuring optimal job scheduling and mitigating GPU memory leaks during continuous AI inference. Infrastructure as Code (IaC): I build reproducible infrastructure. I write modular Terraform to provision the base compute layers and use Ansible to automate node-level configurations, ensuring zero configuration drift across the cluster. Observability & CI/CD: You cannot manage what you cannot see. I deploy the Prometheus/Grafana stack specifically tuned for scraping GPU metrics (via dcgm-exporter), allowing proactive monitoring of GPU health and thermal throttling. I also structure CI/CD pipelines to automate the deployment of these AI models seamlessly. I am based in India and ready to support your global infrastructure immediately. Best regards,
₹35,000 INR in 20 days
0.0
0.0

Chicago, United States
Member since Jul 13, 2025
₹12500-37500 INR
₹500000-1000000 INR
$30-250 USD
$57 USD / hour
$30-250 USD
$15-25 USD / hour
£250-750 GBP
$25-50 USD / hour
$18.5 USD / hour
$50 USD / hour
₹1500-12500 INR
$78 USD / hour
$15 USD / hour
$18.5 USD / hour
₹12500-37500 INR
₹600-1500 INR
₹12500-37500 INR
€8-30 EUR
$2-8 USD / hour
$43 USD / hour