Production Support Specialist
Capco · India
Job Description
Role: Production Support Engineer
Location: Hyderabad
Experience 7-12 years
Apply here:
About the role
We’re looking for an Operations / Production Engineer to keep business-critical services stable, secure and available day to day. You’ll combine strong Linux/Unix fundamentals, practical SQL skills and structured incident management to diagnose issues, restore services quickly and improve long-term reliability.
This role suits someone who enjoys solving real-world production problems, working calmly under pressure and turning recurring incidents into automation, monitoring and preventative improvements.
Key responsibilities
Production support and reliability
- Monitor the health, performance and availability of production services.
- Investigate and resolve incidents, outages, performance degradation and service alerts.
- Perform structured triage, identify business impact and prioritise response accordingly.
- Use Linux/Unix tools to diagnose system, process, network, memory and storage issues.
- Analyse application and system logs to identify root causes and contributing factors.
- Execute approved operational procedures, recovery activities and service restoration plans.
- Participate in on-call or out-of-hours support arrangements, where required.
Monitoring and automation
- Develop and improve monitoring, alerting and service-health checks.
- Reduce manual effort through scripting, automation and repeatable operational tooling.
- Identify recurring incidents and deliver preventative improvements.
- Contribute to capacity, resilience, disaster-recovery and operational-readiness activities.
- Improve runbooks, standard operating procedures and knowledge articles.
Incident and problem management
- Manage incidents from initial report through diagnosis, escalation, resolution and closure.
- Communicate clearly with stakeholders throughout the incident lifecycle.
- Escalate to specialist teams and suppliers when required, providing useful evidence and impact details.
- Support root-cause analysis and post-incident reviews.
- Track corrective and preventative actions through to completion.
- Maintain accurate incident, change and problem records.
Database and data investigation
- Excellent knowledge of SQL queries, joins, aggregation etc…
- Able to identify non-performing SQL and optmise it in coordination with development team
Essential skills and experience
- Experience in Operations Engineering, Production Support, Site Reliability Engineering, Infrastructure Support or a similar role.
- Understanding of incident, change and problem-management practices aligned to IT service-management principles.
- Ability to assess impact, prioritise incidents and work effectively under pressure.
- Strong written and verbal communication skills.
- A disciplined approach to documentation, risk management and operational controls.
- Commitment to security, resilience, service quality and continuous improvement.
- Working experience on PostgreSQL and Linux platform is a must
Desirable skills
- Knowledge of ITIL practices and service-management tooling.
- Familiarity with observability platforms, metrics, dashboards, alerting and distributed tracing.
- Scripting or automation experience using languages such as Bash, Python or PowerShell.
- Experience with deployment pipelines, version control and infrastructure-as-code.
- Understanding of resilience testing, disaster recovery and capacity management.
- Experience working in regulated, financial-services or other highly controlled environments.
Details
| Company | Capco |
| Location | India |
| Type | FULL TIME |
| Niche | marketing |
