- Ensuring Uptime of Critical Systems (Incident Response / Triage)
- Monitor, and Troubleshoot Enterprise Services (Prometheus, Grafana, Splunk)
- Configure Enterprise Services (Ansible, YAML, JSON)
- Must be US Citizen due to government requirement
- Must be able to obtain TS/SCI (active TS is preferred)
- Requires a Bachelor's degree in a STEM field and 5+ years of job-related experience, or a Master's degree plus 3 years of job-related experience.
- Experience monitoring large scale systems and using automation to triage emerging issues
- Experience with Prometheus (preferred) and/or Grafana and Splunk.
- Experience automating Systems Administration Activities (Bash / Python / Ansible are preferred)
- Experience developing recovery procedures for large systems (Backup and Restore, Blue/Green Deployment)
- Linux experience
- Collaborative team player with experience working on teams with diverse engineering skills
- Mixed job experience involving software engineering, systems administration, and network engineering