The right talent can transform your business—and we make that happen. At Collabera, we go beyond staffing to deliver strategic workforce solutions that drive growth, innovation, and agility. With deep industry expertise, a global talent network, and a people-first approach, we connect you with professionals who don’t just fit the role but elevate your business. Partner with us and build a workforce that powers success.
DevOps Engineer - HPC / EDA / SLURM / Azure
Contract: Rancho Cordova, California, US span>
Salary Range: 90.00 - 95.00 | Per Hour
Job Code: 372153
End Date: 2026-11-05
Days Left: 26 days, 11 hours left
Pay Range: $90-$95/hr
-
Support and administer SLURM-based HPC compute environments, including partition configuration and migration planning using Terraform and Ansible.
-
Author formal Method of Procedure (MOP) documents and runbooks for infrastructure changes and service cutovers.
-
Coordinate cross-functionally with EDA/TD NAND teams, storage teams, and IDAM to deliver coordinated platform changes.
-
Administer Azure EDA user environment utilizing Thinlinc (VNC).
-
Develop, maintain, and extend Ansible playbooks and roles for Linux system setup, authentication, and platform configuration.
-
Ensure multi-version Ansible playbook compatibility across SLES 15.
-
Contribute GitHub pull requests, conduct code reviews, and manage inner-source infrastructure repositories.
-
Drive production environment changes through change management workflows using ServiceNow.
-
Integrate and configure enterprise identity systems including Okta, Active Directory, LDAP, and SSSD for Linux/HPC environments.
-
Audit and reconcile Linux user and group identity data (UID/GID) across multiple directory and authentication domains.
-
Validate authentication methods and access behavior across HPC compute and storage environments.
-
Extend SSSD-based corporate authentication to new compute environments and author corresponding Ansible automation.
-
Assess and implement log management strategies, including evaluation of Splunk integration for HPC system logs.
-
Investigate and remediate operational issues in production Linux services (VNC, AutoFS, Datadog, etc.).
-
Produce technical documentation, architecture diagrams, implementation guides, and end-user instructions in Confluence.
-
5+ years of experience in a DevOps, Platform Engineering, or Linux Systems Engineering role.
-
Hands-on HPC cluster administration experience, including SLURM or equivalent workload managers.
-
Demonstrated experience supporting EDA or scientific computing environments.
-
Strong Ansible & Terraform automation skill with production-grade playbook and role development.
-
Demonstrated usage and understanding of the Azure cloud compute environment.
-
Familiarity with enterprise Linux identity and authentication stacks (SSSD, LDAP, AD, Okta).
-
Experience with NetApp or comparable enterprise storage platforms in HPC contexts.
-
Ability to author formal technical documentation (MOPs, runbooks, architecture diagrams).
-
Strong written and verbal communication skill; capable of coordinating across multiple teams.
-
Experience with SUSE Linux Enterprise Server (SLES) 12 and/or 15 in an enterprise environment.
-
Experience migrating configuration artifacts and binaries to Artifactory.
-
Background in semiconductor, storage, or high-tech manufacturing IT environments
The Company offers the following benefits for this position, subject to applicable eligibility requirements: medical insurance, dental insurance, vision insurance, 401(k) retirement plan, life insurance, long-term disability insurance, short-term disability insurance, paid parking/public transportation, paid time off, paid sick and safe time, hours of paid vacation time, weeks of paid parental leave, and paid holidays annually – as applicable.
Job Requirement
- HPC
- EDA
- Azure
- SLURM
Reach Out to a Recruiter
- Recruiter
- Phone
- PRARTHI MISTRY
- prarthi.mistry@collabera.com