Skip to content
ReferMeAJob
LiveRemoteFull-timeApply by 1 Nov 2026

Cloud Site Reliability Engineer (Design, Implement)

Remote

Experience
5–10 years
Employment
Full-time
Work mode
Remote
Salary
Not disclosed
Deadline
Apply by 1 Nov 2026
Posted
2025-01-16

Required skills

SkillExperienceLevel
Python5+ yearsAdvanced
Ansible5+ yearsAdvanced
Terraform4+ yearsAdvanced
Docker3+ yearsAdvanced
Apache NGINX3+ yearsAdvanced
AWS5+ yearsAdvanced
Azure5+ yearsAdvanced
Linux2+ yearsAdvanced

About the role

As a SRE your job entails, architecting, Implementing and managing heterogeneous & diverse tech stacks spanning multiple datacentres and across various cloud providers. Implement and manage enterprise level software, providing hosting and domain related services to millions of customers across the globe. Your role as a SRE is primarily focussed on helping business and development teams grow, roll out new features to the market with a strong commitment to quality and availability. At the same time, you will be an expert detective, diving into complex escalations involving enterprise level technical challenges, Engineering problems, customer connects and platform growth concerns etc. This role will involve the management of short & long term projects under SLA and adherence to deadlines.


Key Responsibilities:


  • Architect and maintain mission critical global hybrid infrastructure spanning multiple datacenters & cloud providers, leveraging primarily open source technologies.
  • Design next generation scalable systems which are highly available, resilient and capable of handling high volume Internet facing web traffic.
  • Be responsible for downtimes and maintain the product SLA, capacity planning of the systems and overall health & performance of large scale production systems.
  • Participate in weekly 24/7 oncall rotation, solving escalated tickets, resolve outages and debug production issues.
  • Work closely with various stakeholders like Engineering, Monitoring and Operations teams, Noc / Soc, customers & business development teams.
  • Challenge the status quo. Empower development teams by transitioning legacy methodologies, platform & technologies to devops principles, cloud native technologies and newer ecosystems without much friction.
  • Strict adherence to automating routine tasks and scripting, with a low tolerance to manual processes.
  • Needs to be data & metric driven. Develop tools and platforms for better system observability & insights.
  • Writing design decision documentation and is keen on implementing overall production best practices with a strong focus on security & encourage right Devops Workflows.
  • Design, develop, and deploy modular cloud-based systems• Educating teams on the implementation of new cloud technologies and initiatives
  • Develop and maintain cloud solutions in accordance with best practices.


Requirements

  • At least 5+ years of experience with Cloud SRE role (OCI, AWS, GCP) is mandatory.
  • Experience on Configuration Management tools such as Puppet, Ansible, Terraform is MUST.
  • Experience with Container such as Kubernetes, Docker is required.
  • Experience with scripting in Python, Golang to write scripts and automate routine tasks.
  • Proven work experience as a Cloud Engineer or similar role.
  • Experience in Load Balancer such as HAProxy, Nginx, F5, dnsdist, Varnish.
  • Experience in Webservers such as Apache, Nginx.
  • AWS and/or GCP certifications preferred, not a must

Role categories

Similar roles

DE

Remote

NewRemoteFull-timeCI/CDRMSCloud Computing
Not disclosed
Apply Now
AO

Remote

RemoteFull-timeLLM inference enginesPythonLinuxGITAnsible
Not disclosed
Apply Now
CD

Remote

RemoteFull-timePythonShell ScriptingDevOpsJitterbit harmonyCI/CD
Not disclosed
Apply Now
SA

Remote

RemoteContractPowerShellMicrosoft IntuneAzureMicrosoft Windows Server AdministrationMS Office
Not disclosed
Apply Now
SG

Remote

RemoteFull-timeAWSUnified Modeling Language - (UML)Go LangRest APICI/CD
Not disclosed
Apply Now