The right talent can transform your business—and we make that happen. At Collabera, we go beyond staffing to deliver strategic workforce solutions that drive growth, innovation, and agility. With deep industry expertise, a global talent network, and a people-first approach, we connect you with professionals who don’t just fit the role but elevate your business. Partner with us and build a workforce that powers success.
Site Reliability Engineer
Contract: Westbrook, Maine, US span>
Salary Range: 40.00 - 59.00 | Per Hour
Job Code: 371889
End Date: 2026-10-24
Days Left: 29 days, 2 hours left
Job Title: Site Reliability Engineer (REMOTE)
Duration: 06 - 12 Months Contract
100% REMOTE
Pay Range- $40-$60/hr
Job Description:
-
Own the reliability, availability, and performance of hosted diagnostic imaging and telemedicine applications running on AWS.
-
Design, build, and maintain monitoring, alerting, and observability using Datadog and the AWS Console to detect and resolve issues before they affect customers.
-
Manage and optimize core AWS infrastructure (EC2, RDS, ECS, SQS, S3, and related services), applying infrastructure-as-code practices for repeatable, auditable deployments.
-
Build and maintain CI/CD pipelines in Jenkins for automated, low-risk deployment of releases and hotfixes.
-
Lead incident response and root cause analysis for production issues; write and maintain runbooks and postmortems to prevent recurrence.
-
Triage and resolve production escalations, fixing defects and shipping mini-releases with minimal customer disruption.
-
Define and track SLIs/SLOs and error budgets for critical services, and use that data to help prioritize engineering work.
-
Partner with development and product teams to build reliability, security, and scalability into new features from design through deployment.
-
Participate in an on-call rotation, providing timely response to production incidents.
-
3+ years of experience in site reliability engineering, DevOps, or production systems support, ideally for cloud-hosted, customer-facing applications.
-
Hands-on experience with AWS services (EC2, RDS, ECS, SQS, S3) and infrastructure-as-code tools (CloudFormation or Terraform).
-
Experience with CI/CD tooling (Jenkins or similar) and automated deployment pipelines.
-
Proficiency with monitoring and observability platforms (Datadog or similar) and building actionable alerting.
-
Scripting and automation skills (Python, Bash, or similar) to reduce manual toil.
-
Working knowledge of relational databases (Oracle, MySQL, SQL Server) and SQL.
-
Strong troubleshooting and incident-response skills, with the ability to stay calm and methodical under production pressure.
-
Excellent communication skills, both verbal and written, including the ability to translate technical issues to non-technical audiences.
-
Ability to work independently and within cross-functional teams in an agile environment.
-
BA/BS in Computer Science, Engineering, or a related field, or equivalent work experience.
-
Experience with containerization and orchestration (Docker, ECS, or Kubernetes).
-
Familiarity with Java/J2EE, .NET, or similar application stacks to support root-cause debugging.
-
Experience handling raw image data or image files, or working in healthcare or other regulated environments.
-
AWS certification (SysOps Administrator, DevOps Engineer, or Solutions Architect).
-
Knowledge of Angular or React front-end stacks, useful for full-stack incident triage.
The Company offers the following benefits for this position, subject to applicable eligibility requirements: medical insurance, dental insurance, vision insurance, 401(k) retirement plan, life insurance, long-term disability insurance, short-term disability insurance, paid parking/public transportation, paid time off, paid sick and safe time, hours of paid vacation time, weeks of paid parental leave, and paid holidays annually – as applicable.
Job Requirement
- Datadog
- aws
- SRE
- site reliability engineer
Reach Out to a Recruiter
- Recruiter
- Phone
- PRARTHI MISTRY
- prarthi.mistry@collabera.com