Full-Time

Site Reliability Engineer

Huntress Talent

Huntress Talent

No salary listed

New York, NY, USA

In Person

Category
DevOps & Infrastructure (1)
Required Skills
Bash
Microsoft Azure
Python
Data Structures & Algorithms
Groovy
Java
AWS
Ansible
DevOps
Google Cloud Platform
Requirements
  • Bachelor’s degree in Computer Science, related technical field or equivalent practical experience
  • 3+ years of experience with system design, algorithms, data structures, analysis, and software design
  • Deep knowledge in DevOps tools and practices, Enterprise standards and security
  • Knowledge of scripting languages (e.g. shellscript, groovy) and main concepts of Object-Oriented programming (Python or Java is a plus)
  • Solid knowledge on Ansible and CI/CD
  • At least 2 years’ experience with AWS, Azure, GCP cloud technologies (S3 bucket administration; EC2; ELB; Security Groups; IGW; ACLs; etc.)
  • At least 3 years of SRE experience
Responsibilities
  • Show ownership of customer success with the platform management.
  • Partner with Delivery, Engineering, and Product to steer SRE alignment and strategy to ensure reliability of the platform deployments
  • Respond to client reliability concerns and agile problem resolution.
  • Lead teams that design, code, test, and deliver software to ensure application performance and resiliency
  • Ability to communicate with various customer teams and navigate them with WF best practices (IT, DevOps, Security, Tech).
  • Strives for environment management automation either by coding it or by leading and influencing developers to build systems that are easy to run in production.
  • Proactively work on the efficiency and capacity planning to set clear requirements and reduce the system resources usage to make cheaper to run for all our customers.
  • Identify parts of the system that do not scale, provides immediate palliative measures, and drives long term resolution of these incidents.
  • Identify Service Level Indicators (SLIs) that will align the team to meet the availability and latency objectives.
  • Proposes and drives architectural changes that affect the whole company to solve scaling and performance problems
  • Measure the risk of introduced features to plan and improve the infrastructure

Company Size

N/A

Company Stage

N/A

Total Funding

N/A

Headquarters

N/A

Founded

N/A