Production Support Engineer
Solvd
Argentina
•50 minutos atrás
•Nenhuma candidatura
Sobre
- Solvd Inc. is a rapidly growing AI-native consulting and technology services firm delivering enterprise transformation across cloud, data, software engineering, and artificial intelligence. We work with industry-leading organizations to design, build, and operationalize technology solutions that drive measurable business outcomes.
- Following the acquisition of Tooploox, a premier AI and product development company, Solvd now offers true end-to-end delivery—from strategic advisory and solution design to custom AI development and enterprise-scale implementation. Our capability centers combine deep technical expertise, proven delivery methodologies, and sector-specific knowledge to address complex business challenges quickly and effectively.
- We are looking for a Production Support Engineer to join a small, high-trust team responsible for the health and reliability of a revenue-critical sales platform. You'll sit between end users — partners and call centers — and engineering, triaging incidents, managing communication, and keeping stakeholders informed and calm when things go wrong.
- You'll own the incident management response process end-to-end — from first alert through post-mortem and corrective action follow-up. This is not a developer role. The technical bar is deliberately calibrated: you need to understand how APIs work, read logs, and interpret what you're seeing — not implement fixes. What matters equally is your ability to translate technical issues into plain language and manage expectations across very different audiences.
- Longevity and genuine interest in the role matter here. This team values people who want to grow with it, not move through it.
- What you'll do
- Monitor platform health and triage incoming incidents — distinguishing critical issues (outages, service degradation) from non-critical ones (bugs, defects).
- Investigate incidents using logging tools — reading API calls, responses, and log data to understand what happened and where.
- Own the incident management response process, post-mortems, and corrective action follow-up.
- Notify stakeholders of critical issues proactively — specifying SLA risk and communicating clearly on status via email, phone, or ticket system.
- Cross-reference tickets across multiple systems and follow defects through the full lifecycle until closure.
- Communicate clearly with partners and call center teams — translating technical findings into plain language and managing expectations throughout resolution.
- Manage the incident queue in Jira and prioritize bugs within engineering sprint cycles.
- Participate in weekly cross-functional meetings with engineering and account/call center management.
- Provide suggestions for continual improvement of applications and processes.
- Join on-call rotations after ramp-up — responding to alerts via OpsGenie within defined SLA windows.
- Basic qualifications
- 2+ years of troubleshooting and resolving issues for applications, servers, or infrastructure environments.
- 2+ years of providing clear status updates on tasks, issues, and resolutions to stakeholders at multiple levels.
- Working knowledge of how APIs function — able to read and interpret API calls and responses; experience with Postman or similar API testing tools.
- Ability to navigate logging and observability tools such as Splunk, Datadog, or Sumo Logic.
- Experience with SQL queries for troubleshooting and ad hoc reporting.
- Basic comfort reading HTML and JSON, and using browser developer tools for investigation.
- Ability to participate in technical bridge calls and follow incidents through to resolution.
- Exceptional communication skills — able to code-switch between technical and non-technical audiences fluidly; this is the hardest skill to train and the most important one for this role.
- Strong time management, prioritization, and organizational skills under pressure.
- Customer service mindset — genuine interest in supporting end users and resolving issues, not just closing tickets.
- Empathy, humility, and comfort with ambiguity — able to investigate complex issues without a clear playbook.
- Available during U.S. Eastern business hours (9 AM – 6 PM ET); Eastern timezone strongly preferred for onboarding and on-call coordination.
- Bachelor's degree in a related field or equivalent work experience.
- Preferred qualifications
- Experience with AWS — Cloud Practitioner level or above.
- Familiarity with Git in a team environment.
- Understanding of infrastructure-as-code concepts — Terraform or similar.
- Familiarity with OpsGenie or similar alerting platforms.
- Experience using AI tooling to amplify troubleshooting and investigation workflows.
- Understanding of engineering deployment lifecycle and release processes.
- Experience with on-call rotation structures and incident severity frameworks.
- Prior exposure to partner or call center communication management during live incidents.
- Experience in a travel, hospitality, or high-volume transactional platform environment.
- What to expect when you join
- Comprehensive onboarding documentation and a structured 6-month ramp to full self-sufficiency.
- On-call rotations begin only when you're ready, with manager backup during early rotations.
- Active alert window is 8 AM–1 AM Eastern; overnight suppression windows are built in.
- SEV-1 incidents are rare — roughly once per quarter or less; the majority of the work happens during business hours.
- When you join Solvd, you'll…
- Shape real-world AI-driven projects across key industries, working with clients from startup innovation to enterprise transformation.
- Be part of a global team with equal opportunities for collaboration across continents and cultures.
- Thrive in an inclusive environment that prioritizes continuous learning, innovation, and ethical AI standards.
- Ready to make an impact?
- If you're excited to build things that matter, champion responsible AI, and grow with some of the industry’s sharpest minds. Apply today and let’s innovate together.
- Solvd is an equal opportunity employer.




