SRE Engineer

SRE Engineer

SRE Engineer

Jobgether

14 minutos atrás

Nenhuma candidatura

Sobre

  • This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Engineer based in Brazil.
  • As an SRE Engineer, you will play a key role in ensuring the reliability, scalability, performance, and resilience of critical digital environments.
  • You will work across cloud infrastructure, Kubernetes, observability, automation, and incident management to keep systems stable and highly available.
  • The role combines proactive engineering with hands-on troubleshooting, helping teams identify risks before they become production issues.
  • You will establish and monitor reliability metrics such as SLIs, SLOs, SLAs, MTTR, and MTTD to continuously improve operational performance.
  • You’ll collaborate closely with multidisciplinary teams to embed reliability practices throughout the software development lifecycle.
  • Automation and Infrastructure as Code will be central to reducing manual work and creating more efficient, consistent operations.
  • This is an opportunity to contribute to a culture of continuous improvement while working on modern, cloud-based and distributed technology environments.
  • n

Accountabilities

  • Define, monitor, and continuously improve reliability indicators, including SLIs, SLOs, SLAs, MTTR, MTTD, and error budgets.
  • Implement and evolve observability, monitoring, alerting, and APM solutions across applications and infrastructure.
  • Monitor latency, traffic, errors, saturation, availability, and overall system performance.
  • Prevent, investigate, and resolve incidents, minimizing their impact on users and business operations.
  • Conduct root cause analyses and define corrective and preventive actions to avoid recurring incidents.
  • Identify operational risks, bottlenecks, single points of failure, and opportunities to strengthen system resilience.
  • Support the design and evolution of highly available, scalable, resilient, and fault-tolerant solutions.
  • Automate operational activities and reduce repetitive manual tasks through scripting, automation, and Infrastructure as Code.
  • Operate and continuously improve Kubernetes and Docker environments.
  • Support capacity planning, business continuity, disaster recovery, and cloud cost optimization initiatives.
  • Participate in deployments and contribute to application stabilization following releases.
  • Collaborate with engineering, development, infrastructure, and other technical teams to incorporate reliability from the earliest stages of solution design.
  • Create and maintain operational dashboards, alerts, procedures, runbooks, and technical documentation.
  • Promote a culture centered on reliability, observability, automation, prevention, and continuous improvement.

Requirements

  • Proven professional experience as a Site Reliability Engineer, SRE, or in an equivalent reliability/platform engineering role.
  • Practical experience with cloud environments, using one or more of GCP, AWS, or Azure.
  • Hands-on knowledge of Kubernetes and Docker.
  • Experience implementing and managing observability, monitoring, alerting, and APM solutions.
  • Strong understanding of SRE concepts and metrics, including SLI, SLO, SLA, MTTR, MTTD, and error budgets.
  • Experience managing, investigating, troubleshooting, and resolving production incidents.
  • Knowledge of application and infrastructure troubleshooting in complex environments.
  • Experience administering Linux environments.
  • Understanding of networking, security, performance, scalability, and high availability.
  • Experience with automation and Infrastructure as Code practices.
  • Hands-on experience with CI/CD pipelines and modern software delivery practices.
  • Strong communication and collaboration skills, with the ability to work effectively across multidisciplinary teams.
  • Analytical, proactive, collaborative, and prevention-oriented mindset.
  • Experience with GKE, EKS, or AKS is a plus.
  • Knowledge of Dynatrace, Datadog, Grafana, Prometheus, or comparable observability platforms is a plus.
  • Experience with ELK Stack, Elasticsearch, and Kibana is desirable.
  • Knowledge of Terraform and Ansible is desirable.
  • Experience supporting critical systems and distributed architectures is an advantage.
  • Experience in financial institutions or other regulated environments is a plus.
  • Experience with cloud capacity management and cost optimization is desirable.
  • Knowledge of disaster recovery and business continuity practices is beneficial.
  • Cloud, Kubernetes, or SRE certifications are considered a plus.

Benefits

  • Meal and food allowance.
  • Home office allowance.
  • Medical insurance.
  • Dental insurance.
  • Life insurance.
  • Birthday Day Off.
  • TotalPass / Wellhub access.
  • Health and wellness support through the Boon Saúde app.
  • Discounts and partnerships with a variety of establishments.
  • Partnerships with educational institutions and other services.
  • Welcome kit.
  • Structured onboarding program.
  • Access to continuous learning and professional development initiatives.
  • Dedicated learning and knowledge-sharing programs.
  • Employee support and engagement initiatives.
  • Fully remote work opportunity.
  • n

How Jobgether works

  • We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
  • We appreciate your interest and wish you the best!
  • Why Apply Through Jobgether?
  • Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
  • #LI-CL1