Senior L1/L2 Support Engineer

Job Description

We are seeking a motivated and experienced Senior L1/L2 Support Engineer to join our Technical Operations team. The successful candidate will provide operational and technical support for digital platforms, applications, and customer-facing systems.

You will serve as a key point of contact for production incidents, service requests, and operational issues, ensuring timely resolution and high system availability. This role requires strong troubleshooting skills, experience in cloud operations, application support, monitoring and observability, ITSM processes, and log analysis.

The successful candidate will work closely with application development, infrastructure, engineering teams, external vendors, and business stakeholders to maintain stable, secure, and reliable services in a fast-paced operational environment.

Roles & Responsibilities

Production Support & Incident Management

  • Provide L1/L2 technical and application support for digital platforms and customer-facing applications.

  • Monitor and manage production incidents, service requests, alerts, and support tickets.

  • Perform incident triage, troubleshooting, escalation, and resolution within agreed service levels.

  • Assess the impact of production incidents and coordinate with relevant technical and business stakeholders.

  • Support major incident management activities and provide timely updates during critical service disruptions.

  • Conduct post-incident reviews and contribute to root cause analysis to prevent recurring incidents.

  • Develop and maintain operational runbooks, support procedures, and knowledge base documentation.

System Monitoring & Reliability

  • Monitor application, infrastructure, and service health using monitoring and observability tools.

  • Analyse system performance, availability, error trends, and capacity utilisation.

  • Configure and fine-tune alerts to improve operational visibility and reduce unnecessary alert noise.

  • Work closely with engineering and infrastructure teams to improve system reliability and operational resilience.

  • Support continuous improvement initiatives related to monitoring, incident reduction, and service availability.

AWS Cloud & Infrastructure Support

  • Provide operational support for applications and infrastructure hosted on AWS.

  • Perform first-level troubleshooting for AWS services, including EC2, ECS, Lambda, networking, load balancers, CDN, and storage services.

  • Investigate infrastructure-related issues affecting application performance and availability.

  • Support deployment verification and post-release monitoring activities.

Application & Technical Support

  • Troubleshoot application issues across web applications, mobile applications, APIs, middleware, and microservices.

  • Analyse application logs, monitoring data, and system traces to identify issues and determine root causes.

  • Work closely with L3 support and engineering teams to reproduce issues and validate fixes.

  • Support integrations with external systems, partners, and enterprise platforms.

  • Participate in application release, deployment, and post-implementation support activities when required.

ITSM & Operational Support

  • Manage incidents, service requests, problems, and change activities in accordance with established ITSM processes.

  • Use ITSM and ticketing platforms such as ServiceNow and Jira Service Management.

  • Take ownership of support tickets from initial investigation through resolution or appropriate escalation.

  • Identify recurring incidents and operational inefficiencies and recommend improvements.

  • Support automation initiatives to reduce repetitive manual support activities.

Job Requirements

  • Minimum 5 years of experience in Application Support, Production Support, Technical Support, IT Operations, Technical Operations, or related roles.

  • Strong experience in troubleshooting and supporting production applications and customer-facing systems.

  • Hands-on experience with AWS cloud services and operational support.

  • Experience with monitoring and observability tools such as Datadog, Dynatrace, New Relic, AppDynamics, Grafana, or AWS CloudWatch.

  • Experience with log analysis tools such as Splunk, Elasticsearch, Kibana, or CloudWatch Logs.

  • Good understanding of APIs, microservices, web services, system integrations, and authentication and authorisation flows.

  • Experience with ITSM processes, including Incident Management, Problem Management, and Change Management.

  • Experience using ticketing platforms such as ServiceNow or Jira Service Management.

  • Strong troubleshooting, analytical, problem-solving, and root cause analysis skills.

  • Ability to work independently and collaboratively in a fast-paced operational environment.

  • Good communication and stakeholder management skills.

  • Ability to participate in a 24/7 support and on-call rotation when required.

Preferred Skills

  • Experience supporting aviation, airline, airport, or travel technology systems.

  • Knowledge of Site Reliability Engineering principles and operational best practices.

  • Basic scripting knowledge in Python, Shell, Bash, or PowerShell.

  • Experience with API Gateway and event-driven architectures.

  • AWS or ITIL certifications will be advantageous.

  • Ability to communicate in Mandarin to support Mandarin-speaking stakeholders and business requirements, where required.

Additional Information

  • Candidates should have strong ownership and be comfortable handling incidents from investigation through resolution.

  • The role may require participation in on-call support, late-night, weekend, or public holiday deployment and support activities when necessary.

See also

要針對這個職缺調整履歷嗎?

目前無法檢查您與這個職缺的符合程度;請先將履歷加入個人檔案,下次即可查看。

A new version of freehire is available