Production Support Engineer

Production Support Engineer

Production Support Engineer

Solvd

Argentina

50 minutos atrás

Nenhuma candidatura

Sobre

  • Solvd Inc. is a rapidly growing AI-native consulting and technology services firm delivering enterprise transformation across cloud, data, software engineering, and artificial intelligence. We work with industry-leading organizations to design, build, and operationalize technology solutions that drive measurable business outcomes.
  • Following the acquisition of Tooploox, a premier AI and product development company, Solvd now offers true end-to-end delivery—from strategic advisory and solution design to custom AI development and enterprise-scale implementation. Our capability centers combine deep technical expertise, proven delivery methodologies, and sector-specific knowledge to address complex business challenges quickly and effectively.
  • We are looking for a Production Support Engineer to join a small, high-trust team responsible for the health and reliability of a revenue-critical sales platform. You'll sit between end users — partners and call centers — and engineering, triaging incidents, managing communication, and keeping stakeholders informed and calm when things go wrong.
  • You'll own the incident management response process end-to-end — from first alert through post-mortem and corrective action follow-up. This is not a developer role. The technical bar is deliberately calibrated: you need to understand how APIs work, read logs, and interpret what you're seeing — not implement fixes. What matters equally is your ability to translate technical issues into plain language and manage expectations across very different audiences.
  • Longevity and genuine interest in the role matter here. This team values people who want to grow with it, not move through it.
  • What you'll do
  • Monitor platform health and triage incoming incidents — distinguishing critical issues (outages, service degradation) from non-critical ones (bugs, defects).
  • Investigate incidents using logging tools — reading API calls, responses, and log data to understand what happened and where.
  • Own the incident management response process, post-mortems, and corrective action follow-up.
  • Notify stakeholders of critical issues proactively — specifying SLA risk and communicating clearly on status via email, phone, or ticket system.
  • Cross-reference tickets across multiple systems and follow defects through the full lifecycle until closure.
  • Communicate clearly with partners and call center teams — translating technical findings into plain language and managing expectations throughout resolution.
  • Manage the incident queue in Jira and prioritize bugs within engineering sprint cycles.
  • Participate in weekly cross-functional meetings with engineering and account/call center management.
  • Provide suggestions for continual improvement of applications and processes.
  • Join on-call rotations after ramp-up — responding to alerts via OpsGenie within defined SLA windows.
  • Basic qualifications
  • 2+ years of troubleshooting and resolving issues for applications, servers, or infrastructure environments.
  • 2+ years of providing clear status updates on tasks, issues, and resolutions to stakeholders at multiple levels.
  • Working knowledge of how APIs function — able to read and interpret API calls and responses; experience with Postman or similar API testing tools.
  • Ability to navigate logging and observability tools such as Splunk, Datadog, or Sumo Logic.
  • Experience with SQL queries for troubleshooting and ad hoc reporting.
  • Basic comfort reading HTML and JSON, and using browser developer tools for investigation.
  • Ability to participate in technical bridge calls and follow incidents through to resolution.
  • Exceptional communication skills — able to code-switch between technical and non-technical audiences fluidly; this is the hardest skill to train and the most important one for this role.
  • Strong time management, prioritization, and organizational skills under pressure.
  • Customer service mindset — genuine interest in supporting end users and resolving issues, not just closing tickets.
  • Empathy, humility, and comfort with ambiguity — able to investigate complex issues without a clear playbook.
  • Available during U.S. Eastern business hours (9 AM – 6 PM ET); Eastern timezone strongly preferred for onboarding and on-call coordination.
  • Bachelor's degree in a related field or equivalent work experience.
  • Preferred qualifications
  • Experience with AWS — Cloud Practitioner level or above.
  • Familiarity with Git in a team environment.
  • Understanding of infrastructure-as-code concepts — Terraform or similar.
  • Familiarity with OpsGenie or similar alerting platforms.
  • Experience using AI tooling to amplify troubleshooting and investigation workflows.
  • Understanding of engineering deployment lifecycle and release processes.
  • Experience with on-call rotation structures and incident severity frameworks.
  • Prior exposure to partner or call center communication management during live incidents.
  • Experience in a travel, hospitality, or high-volume transactional platform environment.
  • What to expect when you join
  • Comprehensive onboarding documentation and a structured 6-month ramp to full self-sufficiency.
  • On-call rotations begin only when you're ready, with manager backup during early rotations.
  • Active alert window is 8 AM–1 AM Eastern; overnight suppression windows are built in.
  • SEV-1 incidents are rare — roughly once per quarter or less; the majority of the work happens during business hours.
  • When you join Solvd, you'll…
  • Shape real-world AI-driven projects across key industries, working with clients from startup innovation to enterprise transformation.
  • Be part of a global team with equal opportunities for collaboration across continents and cultures.
  • Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.
  • Ready to make an impact?
  • If you're excited to build things that matter, champion responsible AI, and grow with some of the industry’s sharpest minds. Apply today and let’s innovate together.
  • Solvd is an equal opportunity employer.