Job Description
Systems Engineer
Job Location:  Singapore
Location Flexibility:  Primary Location Only
Req Id:  11076
Posting Start Date:  8/7/26

Responsibilites: 

  • Manage the day-to-day operations of the HPE Cray EX supercomputing environment, ensuring high availability, stability, performance, and reliability of HPC services. 
  • Administer and maintain HPE Cluster Manager (HPCM) for cluster provisioning, monitoring, health management, software deployment, and lifecycle management. 
  • Manage AMD-based HPE Cray EX compute infrastructure delivering up to 10 PFLOPS of computational performance, ensuring optimal resource utilization and system efficiency. 
  • Administer and optimize HPE ClusterStor Lustre parallel file system with over 10 PB of storage capacity, ensuring high-performance I/O, data integrity, and storage availability. 
  • Manage IBM Storage Scale (formerly GPFS) parallel file system with over 15 PB of storage capacity, including performance tuning, capacity planning, and filesystem maintenance. 
  • Configure, administer, and maintain the PBS Professional workload manager, including queue configuration, scheduling policies, fair-share management, resource allocation, and job troubleshooting. 
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems. 
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability. 
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations. 
  • Manage and troubleshoot the HPE Slingshot high-speed, low-latency interconnect fabric to ensure efficient communication between compute nodes and storage systems. 
  • Monitor overall cluster health, identify performance bottlenecks, perform root cause analysis, and implement corrective and preventive actions to maximize system availability. 
  • Collaborate with infrastructure, storage, networking, and application teams to support HPC platform deployments, upgrades, maintenance activities, and production operations. 
  • Provide technical support to a diverse community of researchers, scientists, engineers, and academic users by troubleshooting application, storage, scheduler, and system-related issues. 
  • Assist users in optimizing HPC applications through performance analysis, job scheduling best practices, parallel computing techniques, and efficient resource utilization. 
  • Conduct user onboarding sessions, technical workshops, and training programs on HPC environment usage, job submission, parallel file systems, and cluster best practices. 
  • Perform software installation, upgrades, patch management, and validation for HPC operating systems, middleware, compilers, MPI libraries, and scientific applications. 
  • Develop and maintain automation scripts using Shell, Python, or similar scripting languages to streamline system administration, monitoring, reporting, and operational tasks. 
  • Maintain comprehensive operational documentation, standard operating procedures (SOPs), architecture diagrams, and technical knowledge base articles. 
  • Participate in incident response, planned maintenance activities, disaster recovery exercises, and root cause analysis to ensure continuous improvement of HPC infrastructure. 
  • Ensure adherence to security policies, operational standards, and best practices while maintaining a secure and highly available HPC environment. 
  • Continuously evaluate emerging HPC technologies and recommend improvements to enhance system performance, scalability, reliability, and operational efficiency.

 

Requirements:

  • 5–7 years of hands-on experience administering High Performance Computing (HPC) environments in enterprise, research, or academic organizations. 
  • Familiarity with configuration management and automation tools such as Ansible, xCAT, Bright Cluster Manager, or Infrastructure-as-Code solutions is an advantage. 
  • Understanding of GPU computing technologies (NVIDIA CUDA, AMD ROCm) and accelerator-based HPC environments is an added advantage. 
Relocation Supported:  No
Visa Sponsorship Approved:  No